Voice AI Agents — a pipeline that listens, looks things up and speaks
Every metric from one trace — quality, failure rate, SLO, gates
Goal
Record spans for each turn of the pipeline, compute latency distributions, WER, response quality, and failure rate from the same raw data, judge them against an SLO, and then build a deployment gate that blocks regressions.
Why it matters
If you look only at average latency and error rate, you cannot see "wrong answers that pass the grounding check" or "errors that were quietly handled well." Metrics must be computed from the same raw data for the numbers to agree, and a gate can be trusted only if you have fed it what it should block. This lab uses the baseline pipeline voicekit.pipeline.Pipeline — the same line as modules 5 and 8 (finalization → rewriting and retrieval → threshold → JSON answer → number check → extraction → TTS) — which reports a span for each stage through run(wav, qid, context, span=기록함수) (the placeholder is the recording function), and lets you deliberately fail a stage with faults={...}. The grader recomputes all summary numbers from your span records and cross-checks them.
Steps
- Leave the 12 queries as 12 traces (a root turn + stage spans) in
/root/voice/eval/spans.jsonl. - Write the p50 and p95 of each stage's span length to
/root/voice/eval/stages.json. - Write the per-query and overall WER, computed from the text of the asr span and the reference text, to
/root/voice/eval/wer.json. - Write factual accuracy, refusal accuracy, and the grounding rate, computed against the answer key, to
/root/voice/eval/quality.json. - Create
/root/voice/eval/spans_faults.jsonl, which reruns with one query each given an empty ASR result, an LLM timeout, and a TTS failure, and/root/voice/eval/failures.json, which counts error turns and user failures. - Set the SLO objectives and create
/root/voice/eval/slo.json, which compares them with the measurements. - Create
/root/voice/eval/gate.py, which compares the baseline and current metrics and exits with code 1 on a regression. - Create this version's metrics
/root/voice/eval/metrics.jsonand/root/voice/eval/report.json, where you ran the gate against the baseline version (/opt/lab/fixtures/voice/eval/baseline.json).
Notes
- One span line:
{"trace_id": 질의 id, "span_id": …, "parent_id": 뿌리의 span_id(뿌리는 null), "name": …, "start_ms": …, "end_ms": …, "status": "ok"|"error"|…, "attrs": {…}}(the placeholders are the query id and the root's span_id, with null for the root). In the attrs of the root turn, putsay,source, andfirst_audio_ms, and in status put the outcome of the turn (ok · refused · degraded · reprompt · failed). - Answer key:
answerableandfactsin/opt/lab/fixtures/voice/rag/queries.jsonl; reference text:/opt/lab/fixtures/voice/queries/refs.tsv. - Percentiles use nearest rank (the ⌈p/100·n⌉-th value after sorting).
- Common mistakes: leaving out of the factual accuracy denominator the answerable questions that were refused (the denominator is all answerable questions), counting every turn with an error as a user failure, and setting the gate's latency tolerance as an absolute value.
- Documents: OpenTelemetry — Traces · Google SRE Book — Service Level Objectives
One turn = one trace
After voice-llm up, run the 12 queries (in the order of /opt/lab/fixtures/voice/rag/queries.jsonl, with the context of a follow-up being the finalized text of the previous query) through Pipeline(), collect the stage spans received by the span hook, add for each turn a root span turn (start 0, end = the maximum end of its children, status = the turn outcome, with say, source, mode, first_audio_ms, and errors in attrs), and write them one per line to /root/voice/eval/spans.jsonl.
The span hook is called as span(name, start_ms, end_ms, status, attrs). Decide the root's span_id first and use it as the parent_id of the children. uuid.uuid4().hex is a simple id.
Per-stage distributions
Group the child spans of spans.jsonl (excluding the root) by name and write n, p50_ms, and p95_ms (nearest rank) of the length (end − start) to /root/voice/eval/stages.json.
The stage with the longest p95 is where users wait most often. Look at the tail of the distribution, not the average.
WER on data that resembles the service
From the attrs.text of the asr span and the original text in /opt/lab/fixtures/voice/queries/refs.tsv, create /root/voice/eval/wer.json with the per-query WER in per_query and the sum of errors ÷ the sum of words in wer (lowercase, punctuation cleanup, keep apostrophes).
You may bring over wer.py from module 3. Compare it with the little over 1% the same model got wrong on LibriSpeech.
Response quality — grounding is not correctness
From the root spans and the answer key, write to /root/voice/eval/quality.json fact_accuracy (with all answerable questions as the denominator, a hit if it did not refuse and say contains one of the facts — case-insensitive), refusal_accuracy (the share of unanswerable questions whose status is refused), and grounded_rate (the share of non-refused answers in which all the numbers in say are in the source document).
Putting the three numbers side by side shows that factual accuracy can be low even when the grounding rate is high. Extract numbers with re.findall(r"\d+(?::\d+)?", …).
Separate errors from user failures
Rerun 01 with Pipeline(faults={"asr_empty": {"q09"}, "llm_timeout": {"q03"}, "tts_error": {"q07"}}) and leave it in /root/voice/eval/spans_faults.jsonl, and write to /root/voice/eval/failures.json turns, error_turns (the number of turns with a child span whose status is error), user_failures (the number of turns whose root status is failed or reprompt), user_failure_rate, and by_type.
In a turn where the LLM died, the system reads out a sentence from the retrieved paragraph (degraded), so the user hears an answer — it is an error, but not a failure. If TTS dies, the user hears silence.
Judge against the SLO
Create /root/voice/eval/slo.json with objectives set to p95_first_audio_ms, fact_accuracy_min (0.5 or higher), and user_failure_rate_max (0.1 or lower), measured holding the first_audio_ms p95 of the spans.jsonl roots, the fact_accuracy of quality.json, and the user_failure_rate of failures.json, and pass holding whether each of the three passed.
Write the objectives in terms of what the user experiences. The failure rate comes from the run where you injected failures — the point of this step is to see whether the injected failures break the objectives. Even if it comes out as a miss, grading looks only at whether the verdict is correct.
A gate that blocks regressions
Create /root/voice/eval/gate.py, invoked as python3 gate.py 기준.json 현재.json (the placeholders are the baseline file and the current file). If wer rises by more than 0.02, fact_accuracy falls by more than 0.10, p95_first_audio_ms rises by more than 20% of the baseline, or user_failure_rate rises by more than 0.05, it prints the reason and exits with code 1, and otherwise exits with 0. The grader tests both directions with hidden baseline and current pairs.
If you keep a table of (name, the direction in which it gets worse, tolerance, whether it is a ratio) for each metric, the rules are visible at a glance. A metric that got better passes no matter how much better it got.
Put this version through the gate
In /root/voice/eval/metrics.json, write wer (wer.json), fact_accuracy (quality.json), p95_first_audio_ms (measured in slo.json), and user_failure_rate (0.0, since this is the normal run), and create /root/voice/eval/report.json holding the exit code of python3 gate.py /opt/lab/fixtures/voice/eval/baseline.json metrics.json in gate_exit and the output lines in gate_output.
The baseline version holds the metrics of a "previous deployment (fictional)." If the gate blocked, the output says which metric caused it — whether to ship that version is for a person to decide.