Voice AI Agents — a pipeline that listens, looks things up and speaks
When it opens its mouth — SSE streaming, time to first token, KV cache
In one line
The moment a voice assistant opens its mouth is not when the LLM has finished writing the whole answer but when it has finished the first sentence. So you do not receive the response as one block but take it as tokens come out (SSE streaming). The time to the first token (TTFT) is mostly the time to compute the prompt (prefill) and is proportional to the prompt length. If you send the same leading part again, the server reuses the KV cache and nearly eliminates that time — but if even the very first character changes, everything after it is recomputed.
Why this was needed
One turn of LabHub's conversation practice is ASR 0.4 seconds + model 0.6 seconds + TTS 2.3 seconds ≈ 3.3 seconds (measured 2026-09-23, backend/app/lang_talk.py). It is a serial structure that hands off to TTS after receiving the whole model answer, so it grows directly as the answer gets longer — which is why that file groups the answer by sentence count and truncates it in code if it overflows ("if the model gets excited and writes five sentences, it takes 12 seconds"). If you receive it by streaming, you can start TTS the moment the first sentence ends. This module measures when that "first sentence" arrives and what delays it.
How it works
llama-server and SSE. The llama-server of llama.cpp opens a /v1/chat/completions of the same shape as OpenAI's. If you send it with "stream": true, the response arrives as Server-Sent Events — a repetition of data: {JSON} lines and blank lines, with new characters in choices[0].delta.content in each chunk, and the last one is data: [DONE]. The first chunk usually has only the role and empty content. If you count that as the first token, TTFT comes out shorter than it really is. The last chunk carries the timings measured by the server (prompt_n · prompt_ms · predicted_n …).
TTFT = wait + prefill + one first token. The model first computes the entire prompt at once to produce the K·V of each layer (prefill), and then writes one token at a time (decode). Prefill is proportional to the number of prompt tokens. With Qwen2.5-0.5B-Instruct Q4_0 from this image loaded on a 2-core CPU Pod (measured on nuc1), the first token of a 35-token prompt took 127 ms and that of a 258-token prompt 801 ms, and generation afterward ran at about 50 tokens per second. That means that the moment you attach a few documents with RAG, the first token becomes several times slower.
KV cache and leading-part matching. The server holds the K·V of the previous request for each slot (conversation seat). If the leading part of a new request is the same as that, it skips that much and computes only the new tokens (cache_prompt). When the same 258-token prompt was sent a second time, prompt_n became 1 and the first token dropped to 24 ms (measured on nuc1). But the comparison is from the front. If you put a changing line such as "Current time: 14:05" at the very start of the system prompt, every request is recomputed from the beginning. Put what changes at the back and what is fixed at the front.
The server in this image is started with --cache-ram 0. The latest llama-server gathers the KV of past prompts in host memory and revives it when a similar request comes, but its default cap is 8,192 MiB, which can exceed the Pod limit (2 GiB), and if left on, old prompts come back to life even after you empty the slot, blurring the cache experiment (measured: even the first prompt sent had a prompt_n of 1). In production, it is a useful feature when there are several conversations.
First sentence and max_tokens. In voice, a long answer means a long wait. You measure separately the moment the first sentence ends — the moment a space comes after a period or question mark — and cap the answer with max_tokens. The reason to wait for the space is to avoid cutting at '3.5' or 'a.m.' (module 7 covers this further).
Barge-in = dropping the connection. If the user breaks in, the remaining generation is to be discarded. If you close the HTTP connection that was reading the stream, llama-server cancels that job (cancel task in the log). If you only stop reading and keep the connection open, the server generates to the end and uses CPU — in a Pod where 2 cores are shared with TTS, the next turn becomes slower by that much.
What it looks like in the field
Small models do not follow instructions well. Even when the 0.5B model in this image was told "in exactly three sentences," it often answered with one sentence or a numbered list. So the operating side does not simply trust what the model says but constrains length in code — max_tokens, truncating by sentence count, splitting the first sentence short and reading it out first (module 7). LabHub's conversation practice pins the sentence count in the instruction and truncates in code when it overflows for the same reason.
What you will do in the next lab
Start the server with voice-llm up and save /health and /props. Build sse.py, which reads SSE line by line, record the time of each chunk, and measure the first-token latency five times with the cache off. Attach 0, 4, and 12 documents and measure prefill, and see the cache reusing the leading part and one line at the very front breaking the cache. Measure the moment the first sentence ends, and cut the connection after 0.5 seconds to check with /slots and the log whether the server stops.