Voice AI Agents — a pipeline that listens, looks things up and speaks
What to measure before you can trust it — spans, quality metrics, failure rates and gates
In one line
One turn of a voice assistant goes through several stages, and each stage slows down and fails on its own. So you leave behind one trace per turn and one span per stage, and compute every metric — per-stage latency distributions, WER, response quality, failure rate — from that same raw data. The error rate and the failure rate users experience are different numbers, and wrong answers that pass the grounding check show up only against an answer key. Finally, you put in a gate that blocks deployment if things are worse than the baseline version.
Why this was needed
A report like "average response 1.2 seconds, doing well" hides three things. Which stage is slow (distribution and stage), which turns failed (the kind of failure), and whether the answer was right (quality). As seen in the earlier modules, this pipeline produces wrong answers that pass the grounding check (module 5) — even if it answers a Saturday question with weekday hours, the numbers are in the document, so it passes. On a dashboard that measures only latency and errors, this wrong answer is counted as a "success."
How it works
Traces and spans. It follows the shape of OpenTelemetry — each span has a trace_id, span_id, parent_id, name, start, end, status, and attributes. One turn has one root span (turn) and child spans (asr, retrieve, llm, verify, tts). The attributes hold the finalized text, the retrieved documents, the answer, and the source. From this single raw data come the latency distributions (span lengths), WER (the asr span's text versus the reference), quality (the root span's answer versus the answer key), and the failure rate (status). If you keep separate logs for each metric, you end up counting different turns and the numbers do not agree with each other.
WER revisited. On 12 synthesized phone queries, this pipeline's ASR WER was almost 20% (a little over 1% on LibriSpeech in module 3). It is the same model. You can see in numbers why the evaluation data must resemble the service.
Three quality metrics. (1) Factual accuracy — for answerable questions, does the answer contain one of the ground-truth facts ('9 a.m.', '$80' …)? (2) Refusal accuracy — did it refuse the unanswerable questions? (3) Grounding rate — are the numbers in the spoken answer in the cited document? This pipeline had 100% refusal accuracy and 100% grounding rate, but factual accuracy was 50% (measured on the cluster's nuc1 node). With the same model, the same seed, and temperature 0, it was 30% on a Mac (arm64) — when the CPU differs, the order of floating-point computation differs and a few answers of the 0.5B model change. So you measure again in your own run, and build the baseline on the same kind of node. This happens because the 0.5B model copies a different number from the document, or because the extracted sentence that replaced an answer that failed the check misses the question. Grounding is not correctness.
Errors and failures are different. It becomes clear if you deliberately fail three stages. If the ASR produces empty text, the user hears "please say that again" (a re-prompt — a user failure). If the LLM times out but the system reads out a sentence from the retrieved paragraph, the user hears an answer (a degraded response — an error, but not a failure). If TTS fails, the user hears silence (a failure). Of the 3 turns with errors, 2 were user failures. If you keep only the error rate on the dashboard, you cannot tell "errors that were quietly handled well" from "errors that abandoned the user."
SLOs and gates. An SLO (service level objective) is written in terms of what the user experiences — for example, first sound p95 at most 2.5 seconds, factual accuracy at least 0.6, user failure rate at most 5%. You compare it with the measurements and judge pass or miss. A gate goes one step further and looks at whether it got worse than the baseline version. Each metric has a different direction and tolerance — WER and failure rate are bad when they rise (absolute margin), factual accuracy is bad when it falls, and latency wobbles from machine to machine, so you look at it as a ratio (for example, 20%). A gate is also code, so feed it what it should block and test that the red light comes on.
What it looks like in the field
Evaluation data starts small and grows with each incident. This course's answer key is just the ground-truth documents and facts attached by hand to 12 queries. Even so, it is enough to show that factual accuracy is only a little over half, and that a 100% grounding rate hides it. In production, you pull turns that were handed off (handoff) and re-prompt turns from real calls as samples and keep adding them to the answer key. The fact that the LabHub repository writes "whenever you create a new gate, feed it what it is meant to block and check that the red light comes on" (AGENTS.md) is the same principle.
It should also be recorded that latency metrics depend on the machine. The same pipeline takes tens of percent longer or shorter depending on the node's CPU. So at the gate, latency is compared against a baseline measured on the same machine or looked at as a ratio, and "where it was measured" is written alongside the baseline's metrics. If you judge by a single absolute threshold, you get a false regression every day you run on a slow node. Conversely, quality metrics (WER, factual accuracy) do not depend on the machine, so running again with the same input and the same seed should give the same numbers — if they come out different, that is itself an event to investigate.
What you will do in the next lab
You hook span recording into the baseline pipeline and leave 12 queries as traces. From the same raw data you compute per-stage p50/p95, WER, factual and refusal accuracy, and the grounding rate. You fail three stages to separate error turns from user failures, and set and judge an SLO. You build gate.py, which blocks regressions against the baseline version, have it tested against hidden cases, and run the gate with this version's metrics.