Voice AI Agents — a pipeline that listens, looks things up and speaks
The latency budget — from the moment the caller stops talking to the first sound
In one line
The latency a user feels is from the moment they stop speaking to the moment the first sound comes out. It is the sum of the endpoint decision (audio time) + finalization + retrieval + the first LLM piece + the first TTS sound (wall-clock time). Connected in series, the whole LLM and the whole TTS are added; run overlapped, only the time of the first piece is added. You have to write a budget for each stage and compare against a distribution (p50, p95) to see where to cut.
Why this was needed
The numbers measured for each component in the earlier modules are all small when looked at separately — finalization tens of ms, the first token a hundred or so ms, TTS RTF 0.04. But when you join them, it becomes seconds. The 3.3 seconds of LabHub's conversation practice is also the serial sum of three components (ASR 0.4 + model 0.6 + TTS 2.3) (backend/app/lang_talk.py). How you joined them makes a bigger difference than which component is slow.
How it works
Two clocks. While the user is speaking, streaming ASR is already transcribing. So that time is not latency. Latency starts after speech ends. But the moment you know "speech has ended" (the endpoint) is later than the actual end of speech — it is certain only after the VAD has seen the end and it has stayed quiet for another gap (module 2). This share is measured in audio time: endpoint − actual end of speech. On this course's queries, with gap 0.5 seconds, the median was about 0.57 seconds. What happens after the endpoint (finalization, retrieval, LLM, TTS) is measured on the wall clock, with the endpoint moment as 0. You do not mix the two clocks; you add them at the end.
Serial and overlapped. In serial, TTS synthesizes the whole thing after the LLM has written the whole answer — first sound = finalization + retrieval + the whole LLM + the whole TTS. In overlapped, pieces are passed to TTS as they are produced while the LLM writes — first sound = finalization + retrieval + (until the first piece is produced) + (synthesis of the first piece). On five identical queries, the median first sound was 1,814 ms serial and 1,593 ms overlapped (values measured on a 2-CPU Pod on the cluster's nuc1 node — they differ per node, so look at the numbers you measure in the lab). The reason the gain is smaller than expected is that, as seen in module 7, the LLM and TTS share the same 2 cores.
Draw it with a trace. A bar chart reads faster than a table of numbers. The Chrome trace format (Trace Event Format) is a list of events {"name", "ph": "X", "ts", "dur", "tid"} where ts and dur are in microseconds. If you open it in chrome://tracing or Perfetto, you can see the stretch where the LLM bar and the TTS bar overlap. This format leads into the span records of module 9.
A distribution, not an average. The unpleasantness of a conversation comes from the occasional long silence. So you look at p95 along with the median (p50). With 10 samples, p95 is the slowest one (nearest rank). When samples are few, an interpolated percentile produces a value that does not exist, so this course uses nearest rank.
Budget. Working backward from the observation that turn-taking gaps in human conversation are generally a few hundred ms (module 2), you write a budget for each stage — for example, endpoint 700 · finalization 150 · retrieval 100 · first token 400 · first sound 600 ms. The stages where the measured median exceeds the budget are where to cut. Some shares do not shrink with the model (endpoint). Some shrink with the prompt (first token — the cache and a short context). Some shrink with the seams (first sound — overlapping and a short first piece).
Memory is a budget too. This Pod's limit is 2 GiB. The Python process (streaming ASR, TTS, embedding) used up to about 370 MB and the LLM server about 730 MB (measured on nuc1; under a burst of requests the LLM server rose to 830 MB) — together a little over half. The moment you load one more model or grow the LLM to 1.5B, you hit the limit. If you exceed it, the kernel kills the process (OOM), and the user hears an unexplained silence.
What it looks like in the field
LabHub's conversation practice chains one turn as "whole recording → ASR → model → TTS," and accepts only one person at a time so as not to fight over the GPU (lang_talk.py). To cut 3.3 seconds in that structure, streaming input (the machine decides the endpoint), overlapping the model stream with TTS, and a short first sentence come before making components faster. This module measures that effect directly with a small model.
There is one thing to be careful about when measuring. This Pod has a limit of 2 CPU cores, but nproc inside the container reports the node's core count (24). The BLAS that numpy uses starts that many threads and they push each other out within 2 cores — a short np.dot run 3,200 times to find the echo delay went from 0.11 seconds to 52 seconds (measured while building this image). So the image pins OPENBLAS_NUM_THREADS=2. In experiments that measure latency, a single environment difference like this can throw results off by hundreds of times.
What you will do in the next lab
You measure the endpoint latency of 12 queries in audio time. You chain finalization → retrieval → LLM → TTS in series and measure, then measure overlapped with the same queries and compare. You draw one overlapped run in the Chrome trace format and compute the user-perceived latency p50/p95 over 10 queries. You compare the per-stage budget with the measurements, and measure the memory of the two processes and compare it with the Pod limit.