top of page

Senior Voice AI Engineer - Evaluation & Reliability

About the Role

We’re looking for a senior engineer to own the quality, evaluation, and reliability of a voice AI system used by real patients. 


This is not a role for building a speech stack from scratch. We use a third-party voice platform for telephony, speech recognition, synthesis, end-of-speech detection, and interruption handling. 


Your job is to determine whether the full system is working well enough to put in front of patients — and to say no when the evidence says it isn’t. 


What you’ll own


You will own the evaluation and reliability layer around our production voice assistant. 


That includes: 


  • Measuring end-to-end conversational latency, including backend round trips. 

  • Evaluating end-of-speech detection, silence handling, and turn-taking. 

  • Testing barge-in, interruptions, caller corrections, hesitation, and refusal to answer. 

  • Measuring speech-recognition quality under real telephone conditions. 

  • Building and maintaining automated conversation regression tests. 

  • Creating realistic scenario sets based on production calls and known failure modes. 

  • Defining release thresholds and go/no-go criteria. 

  • Investigating failures across the voice platform, backend, prompts, and integrations. 

  • Turning production incidents into permanent regression tests. 


The core question you will own is: 


Did this change make the voice experience better, worse, or unsafe to release? 


What we’re looking for 


Production voice experience 


You have operated a phone or voice assistant used by real users. 


You should be able to explain: 


  • What its end-to-end response latency was. 

  • How you measured it.

  • Which percentiles you monitored. 

  • Where latency accumulated across the system.


Voice tuning experience 


You have tuned voice behaviour in production, including: 


  • End-of-speech detection. 

  • Endpointing and silence thresholds. 

  • Barge-in and interruption handling. 

  • Long caller pauses. 

  • False interruptions. 

  • Callers correcting themselves mid-answer. 


Experience working with a closed or vendor-controlled voice platform is especially valuable. 


Telephone ASR experience 


You understand how speech recognition changes when audio comes through PSTN or VoIP instead of a high-quality microphone. 


More importantly, you have measured it. 


You should be comfortable evaluating WER/CER as well as business-critical recognition errors such as names, dates, numbers, medications, or other important entities. 


Evaluation systems 


This is the centre of the role. 


You have built or owned a systematic way to test conversational systems, such as: 


  • Conversation replay. 

  • Scenario-based test suites. 

  • Simulation. 

  • Golden conversation sets. 

  • Automated regression testing. 

  • Human review combined with automated scoring. 


We care less about the exact framework and more about whether it allowed you to reliably compare one release against another. 


Engineering 


You have strong Python skills and are comfortable building evaluation tools, API integrations, data pipelines, and automated tests.


Tech stack 


You would likely work with: 


  • Python — evaluation harnesses, automation, analysis. 

  • pytest + Pydantic — scenario definitions and regression tests. 

  • Pandas / Polars + NumPy — latency and quality analysis. 

  • jiwer — WER/CER and ASR evaluation. 

  • FFmpeg / SoX / librosa — audio inspection and telephone-audio testing. 

  • PostgreSQL — storing evaluation runs, scores, and results. 

  • S3-compatible storage — recordings, transcripts, and test artifacts. 

  • OpenTelemetry — end-to-end tracing and latency measurement. 

  • Grafana / Metabase — quality and reliability dashboards. 

  • GitHub Actions / GitLab CI — automated regression testing before releases. 

  • Twilio / SIP tooling — controlled phone-call testing. 

  • HTTP APIs — integration with the voice platform and backend services. 


You do not need experience with every tool above. We care more about whether you understand the underlying problem and know how to build reliable evidence around it. 


Metrics you may own 


Typical metrics include:


  • p50 / p95 / p99 response latency. 

  • End-of-speech detection accuracy. 

  • False and late endpoint rates. 

  • Barge-in success rate. 

  • False interruption rate. 

  • WER / CER. 

  • Critical-entity recognition accuracy. 

  • Task completion rate. 

  • Conversation repair success. 

  • Backend and vendor error rates. 

  • Safety-critical failure rates. 


Nice to have


  • SIP, WebRTC, Asterisk, Twilio, or similar telephony experience. 

  • Experience with LLM or conversational-agent evaluation. 

  • Experience handling sensitive or regulated data. 

  • Healthcare or another safety-sensitive domain. 

  • German-language voice systems.


You do not need 


You do not need to: 


  • Speak German. 

  • Build streaming audio infrastructure. 

  • Build STT or TTS systems. 

  • Build a turn-taking engine. 

  • Train models from scratch. 

  • Know PHP or our existing backend. 


Another team owns the backend endpoints. Your focus is the quality of the complete conversation. 


What success looks like 


We want to reach a point where:


  • Every important conversation change is tested before release. 

  • Production failures become reproducible regression tests. 

  • Voice latency is measured objectively. 

  • ASR quality is tested under realistic phone conditions. 

  • Interruption and endpointing changes can be compared quantitatively. 

  • Every release has explicit quality thresholds. 


Most importantly, the company has someone who can independently answer: 


“Is this version good enough to put in front of patients?”


Apply to  marko@apexteamhiring.com

bottom of page