About the Role
We’re looking for a senior engineer to own the quality, evaluation, and reliability of a voice AI system used by real patients.
This is not a role for building a speech stack from scratch. We use a third-party voice platform for telephony, speech recognition, synthesis, end-of-speech detection, and interruption handling.
Your job is to determine whether the full system is working well enough to put in front of patients — and to say no when the evidence says it isn’t.
What you’ll own
You will own the evaluation and reliability layer around our production voice assistant.
That includes:
Measuring end-to-end conversational latency, including backend round trips.
Evaluating end-of-speech detection, silence handling, and turn-taking.
Testing barge-in, interruptions, caller corrections, hesitation, and refusal to answer.
Measuring speech-recognition quality under real telephone conditions.
Building and maintaining automated conversation regression tests.
Creating realistic scenario sets based on production calls and known failure modes.
Defining release thresholds and go/no-go criteria.
Investigating failures across the voice platform, backend, prompts, and integrations.
Turning production incidents into permanent regression tests.
The core question you will own is:
Did this change make the voice experience better, worse, or unsafe to release?
What we’re looking for
Production voice experience
You have operated a phone or voice assistant used by real users.
You should be able to explain:
What its end-to-end response latency was.
How you measured it.
Which percentiles you monitored.
Where latency accumulated across the system.
Voice tuning experience
You have tuned voice behaviour in production, including:
End-of-speech detection.
Endpointing and silence thresholds.
Barge-in and interruption handling.
Long caller pauses.
False interruptions.
Callers correcting themselves mid-answer.
Experience working with a closed or vendor-controlled voice platform is especially valuable.
Telephone ASR experience
You understand how speech recognition changes when audio comes through PSTN or VoIP instead of a high-quality microphone.
More importantly, you have measured it.
You should be comfortable evaluating WER/CER as well as business-critical recognition errors such as names, dates, numbers, medications, or other important entities.
Evaluation systems
This is the centre of the role.
You have built or owned a systematic way to test conversational systems, such as:
Conversation replay.
Scenario-based test suites.
Simulation.
Golden conversation sets.
Automated regression testing.
Human review combined with automated scoring.
We care less about the exact framework and more about whether it allowed you to reliably compare one release against another.
Engineering
You have strong Python skills and are comfortable building evaluation tools, API integrations, data pipelines, and automated tests.
Tech stack
You would likely work with:
Python — evaluation harnesses, automation, analysis.
pytest + Pydantic — scenario definitions and regression tests.
Pandas / Polars + NumPy — latency and quality analysis.
jiwer — WER/CER and ASR evaluation.
FFmpeg / SoX / librosa — audio inspection and telephone-audio testing.
PostgreSQL — storing evaluation runs, scores, and results.
S3-compatible storage — recordings, transcripts, and test artifacts.
OpenTelemetry — end-to-end tracing and latency measurement.
Grafana / Metabase — quality and reliability dashboards.
GitHub Actions / GitLab CI — automated regression testing before releases.
Twilio / SIP tooling — controlled phone-call testing.
HTTP APIs — integration with the voice platform and backend services.
You do not need experience with every tool above. We care more about whether you understand the underlying problem and know how to build reliable evidence around it.
Metrics you may own
Typical metrics include:
p50 / p95 / p99 response latency.
End-of-speech detection accuracy.
False and late endpoint rates.
Barge-in success rate.
False interruption rate.
WER / CER.
Critical-entity recognition accuracy.
Task completion rate.
Conversation repair success.
Backend and vendor error rates.
Safety-critical failure rates.
Nice to have
SIP, WebRTC, Asterisk, Twilio, or similar telephony experience.
Experience with LLM or conversational-agent evaluation.
Experience handling sensitive or regulated data.
Healthcare or another safety-sensitive domain.
German-language voice systems.
You do not need
You do not need to:
Speak German.
Build streaming audio infrastructure.
Build STT or TTS systems.
Build a turn-taking engine.
Train models from scratch.
Know PHP or our existing backend.
Another team owns the backend endpoints. Your focus is the quality of the complete conversation.
What success looks like
We want to reach a point where:
Every important conversation change is tested before release.
Production failures become reproducible regression tests.
Voice latency is measured objectively.
ASR quality is tested under realistic phone conditions.
Interruption and endpointing changes can be compared quantitatively.
Every release has explicit quality thresholds.
Most importantly, the company has someone who can independently answer:
“Is this version good enough to put in front of patients?”
Apply to marko@apexteamhiring.com