Voice AI Agents — a pipeline that listens, looks things up and speaks
From the end of speech to the first sound — measure serial vs overlapped and check against the budget
Goal
Chain listening through speaking in one line and measure the latency the user feels. Compare serial and overlapped with the same queries, and find where to cut using distributions and budgets.
Why it matters
Even if each component is fast, joining them makes seconds. To know where to cut, you have to look at "how they were joined" and "which stage exceeds its budget" as distributions. Instead of absolute times that differ from machine to machine, the grader looks at the order inside the records (finalization ≤ retrieval ≤ first token ≤ end of LLM), the comparison of the two methods measured on the same machine (overlapped < serial), and the percentiles and budget overruns recomputed from the raw data. The endpoint latency (audio time) is cross-checked by the grader running the VAD again.
Steps
- Write to
/root/voice/pipeline/endpoint.jsonthe difference between the endpoint (the last speech end seen by the VAD + gap 0.5 seconds) and the actual speech end (the fourth column ofrefs.tsv) for the 12 queries (/opt/lab/fixtures/voice/queries/q01~q12.wav). - Create
/root/voice/pipeline/pipe.py, which chains finalization → retrieval → LLM (whole) → TTS (whole), and write the records of running q01, q03, q05, q07, and q09 in series to/root/voice/pipeline/serial.jsonl. - Run the same five queries overlapped, passing LLM pieces to TTS as they are produced, and write them to
/root/voice/pipeline/overlap.jsonl. - Draw one overlapped run in the Chrome trace format to
/root/voice/pipeline/trace.json. - Run the 10 answerable queries (q01–q10) overlapped and write the user-perceived latency (endpoint latency + first sound after the decision) and p50 and p95 to
/root/voice/pipeline/e2e.json. - Write the per-stage budget, the measured median, and the stages that exceeded the budget to
/root/voice/pipeline/budget.json. - Write the peak RSS of the Python process, the RSS of the LLM server, and the Pod memory limit to
/root/voice/pipeline/memory.json. - Create
/root/voice/pipeline/report.jsoncollecting the measurements.
Notes
- Times are wall-clock ms with the endpoint set to 0. Feeding all the sound into the streaming ASR (while the user is speaking) is not measured; measure from the trailing silence and
input_finished(). - Record fields:
id,asr_ms,retrieve_ms,llm_first_ms,llm_done_ms,first_audio_ms,tts_done_ms,text— all are cumulative times measured from 0. - Components needed:
voicekit.models(streaming_asr, tts, Embedder),voicekit.llm.chat_stream, the paragraph search from module 5, and the piece cutting from module 7.voice-llm upcomes first. - Common mistakes: mixing audio time and wall clock, writing the first sound near the first token in the serial record, and putting the trace's ts and dur in ms (they are microseconds).
- Documents: Trace Event Format · Perfetto UI · llama.cpp server
Endpoint latency — in audio time
For each of the 12 queries in /opt/lab/fixtures/voice/queries/refs.tsv (name, original text, utterance start, utterance end), write to /root/voice/pipeline/endpoint.json (gap_s, rows) the end of the last segment of the silero VAD (threshold 0.5 · min_silence 0.1 · min_speech 0.1) as vad_end_s, that plus a gap of 0.5 seconds as endpoint_s, and (endpoint_s − 발화 끝) × 1000 as delay_ms (the placeholder is the utterance end).
All values are times inside the sound file — it is a share that does not shrink even if the machine is fast, so you measure it separately. The actual speech end is exact because it is where the synthesized speech was placed.
Chain it in series
In /root/voice/pipeline/pipe.py, create run(wav, mode). After feeding the sound into the streaming ASR 100 ms at a time (not measured), set the endpoint moment to 0 and return the times for: finalization (asr_ms) → paragraph top-3 retrieval with the query minus filler words (retrieve_ms) → the LLM stream with the retrieved paragraphs attached (first piece llm_first_ms, end llm_done_ms, max_tokens 60) → TTS, through to the first sound (first_audio_ms) and the end of synthesis (tts_done_ms). If mode is 'serial', synthesize the whole thing at once after receiving the entire answer. Run q01, q03, q05, q07, and q09 in series, add id and text (the finalized result), and write them to /root/voice/pipeline/serial.jsonl.
Create the models (ASR, TTS, embedding) and the KB embeddings only once, when the module is imported. If you run TTS in a thread, you can use it as is in the next step's overlap. In serial, the first sound is the end of synthesis.
Run overlapped
Make run(wav, 'overlap') cut pieces (a sentence end · before a comma of more than five words · twelve words) as LLM text accumulates and pass them to the TTS thread. Run the same five queries overlapped and write them to /root/voice/pipeline/overlap.jsonl. Grading: the median first sound must be faster than serial.
The comparison is fair only if the one difference between serial and overlapped is "when it is passed to TTS." Separate them with mode inside the same run function.
Draw it as bars
From the first record of overlap.jsonl, create {"traceEvents": [...]} with the bars asr_final, retrieve, and llm (all tid 1) and tts (tid 2, from the first token to the end of synthesis) as "ph": "X" events, and the first sound as a "ph": "i" event, and save it to /root/voice/pipeline/trace.json. ts and dur are in microseconds.
ms × 1000 = µs. If you drag the file into chrome://tracing or ui.perfetto.dev, you can see the LLM and TTS overlapping.
The distribution of the latency the user feels
Run q01–q10 overlapped and, for each query, write endpoint_ms (delay_ms from endpoint.json), first_audio_ms, and e2e_ms (their sum) in rows, and write p50_ms and p95_ms (nearest rank) of e2e_ms to /root/voice/pipeline/e2e.json.
This is where you first add audio time (endpoint latency) and wall-clock time (first sound after the decision). The p95 of 10 samples is the largest value.
Compare against the budget
Create /root/voice/pipeline/budget.json with budget_ms holding the budgets for five stages (endpoint, asr_final, retrieve, llm_first_token, first_audio_after_token) (positive numbers, total at most 2500), measured_p50_ms holding the medians — for endpoint, delay_ms from endpoint.json, and for the others, from overlap.jsonl: asr_ms · (retrieve_ms − asr_ms) · (llm_first_ms − retrieve_ms) · (first_audio_ms − llm_first_ms) — and over_budget holding the names of the stages where the measurement exceeded the budget.
The difference between cumulative times is that stage's share. For each stage that overflowed, think about what could cut it — the model, the prompt, or the seams.
Memory is a budget too
After running one query with pipe, write to /root/voice/pipeline/memory.json the VmHWM (peak RSS) in /proc/self/status of this Python process as python_peak_mb, the VmRSS of the server found with pgrep -f llama-server as llm_rss_mb, their sum as total_mb, and the limit in /sys/fs/cgroup/memory.max as limit_mb (null if it is not a number).
VmHWM and VmRSS come out in kB — divide by 1024 to get MB. The limit is in bytes. What percent of the limit the two processes use is the answer to whether you can load one more model.
Report
In /root/voice/pipeline/report.json, write serial_first_audio_p50_ms and overlap_first_audio_p50_ms (the medians of first_audio_ms in the two jsonl files), e2e_p50_ms and e2e_p95_ms, over_budget, and total_mb.
Compute or copy them from the files of the previous steps.