LabHub
Get started
Learn Learning paths Courses

Voice AI Agents — a pipeline that listens, looks things up and speaks

When does the first token arrive — read SSE by hand and measure prefill, cache and cancellation

Continue in LabHub

Goal

Read a streaming response yourself from a small LLM server inside the Pod, measure first-token latency, prefill, the KV cache, and the first-sentence time, and cut the connection to stop generation.

Why it matters

A voice assistant can start speaking the moment the first sentence ends. The biggest share of what delays that moment is prompt computation (prefill), and the cheapest way to reduce it is the KV cache. You have to look at both as numbers to judge how to compose the prompt. The grader does not query the LLM server; it reads only the raw SSE, measurement JSON, and server logs you saved, and instead of absolute times that differ from machine to machine, it looks at relationships within the same run (first token < total, long prompt > short prompt, cache hit < no cache).

Steps

  1. Start the server with voice-llm up and save /health to /root/voice/llm/health.json and /props to /root/voice/llm/props.json.
  2. Create /root/voice/llm/sse.py, which reads SSE line by line, and save the lines received as they are to /root/voice/llm/sse_raw.txt and the content chunks with their times to /root/voice/llm/chunks.jsonl.
  3. Send the same question five times with cache_prompt: false and write the first token, total time, and token count to /root/voice/llm/ttft.json.
  4. Attach 0 · 4 · 12 KB documents to the system prompt (cache off) and write prompt_n and the first-token time to /root/voice/llm/prefill.json.
  5. Turn the cache on and send the same long prompt twice, and once more with one line containing the time at the very front, and write it to /root/voice/llm/cache.json.
  6. Receive a three-sentence answer by streaming and write the time the first sentence ended to /root/voice/llm/sentence.json.
  7. Request a long answer, cut the connection after 0.5 seconds, and write the time until the server stops to /root/voice/llm/cancel.json.
  8. Create /root/voice/llm/report.json collecting the measurements.

Notes

Start the server and see what came up

Start the LLM server with voice-llm up, then save curl -s localhost:8080/health to /root/voice/llm/health.json and curl -s localhost:8080/props to /root/voice/llm/props.json.

If /health returns {"status":"ok"}, the model has finished loading. /props contains the model path, the context length (default_generation_settings.n_ctx), and the number of slots (total_slots).

Read SSE line by line

In /root/voice/llm/sse.py, create stream(messages, max_tokens=64, cache_prompt=True, raw_out=None) that sends to /v1/chat/completions with stream: true, writes the lines received as they are to raw_out, collects {"t_ms": 요청 직전부터의 ms, "text": 조각} (the placeholders are the ms since just before the request and the chunk) for every chunk that has content, and returns (the list of chunks, total ms, timings). Run it once with the system prompt 'You are a clinic phone assistant. Answer in one short sentence.' and the question 'What should a new patient bring?' to create /root/voice/llm/sse_raw.txt and /root/voice/llm/chunks.jsonl.

If a line starts with 'data: ', parse what follows as JSON. 'data: [DONE]' is the end. Skip chunks where choices[0].delta.content is empty (the role in the first chunk). Measure time with time.perf_counter().

Measure first-token latency five times

Send the question from 02 five times with cache_prompt=False, and create /root/voice/llm/ttft.json with ttft_ms (the time of the first content chunk), total_ms, and tokens (predicted_n in timings) for each run in runs, and the medians of the two times in ttft_p50 and total_p50.

With the cache off, the entire prompt is recomputed every time. The median is the third of the five sorted values. The first run may be a little slow because the server has just started — that is why you measure several times rather than once.

A long prompt delays the first token

Read /opt/lab/fixtures/voice/kb/*.md in name order, attach the first 0 · 4 · 12 of them to the system prompt, send the question 'When is the clinic open on Saturday?' with cache_prompt=False and max_tokens=32, and write the three results {"docs": n, "prompt_n": …, "prompt_ms": …, "ttft_ms": …} as an array in /root/voice/llm/prefill.json.

prompt_n and prompt_ms are in the timings of the last chunk. Notice that the first token gets delayed almost in proportion as tokens increase. If the cache is on, the leading part that overlaps with the previous request drops out and the numbers blur.

KV cache — the same leading part is skipped

With a long system prompt containing all the KB documents, first send one different short request ('hi') to empty the slot, then send it twice with cache_prompt=True (first and second), and once more with 'Current time: 14:05. ' attached at the very front of the system prompt (changed_prefix), and write the prompt_n and ttft_ms of each to /root/voice/llm/cache.json.

The server compares the K·V of the previous request left in the slot with the new request from the front and skips the same amount. If one line at the very front differs, it differs from the first token, so nothing can be reused. Also think about what happens if you attach the changing line at the very end.

When does the first sentence end

Stream the system prompt 'You are a clinic phone assistant. Answer in exactly three short sentences.' and the question 'How do I prepare for a fasting blood test?' with max_tokens=120, and create /root/voice/llm/sentence.json (ttft_ms, first_sentence_ms, total_ms, first_sentence, text) with the time of the first chunk after which a space follows [.!?] in the concatenated text as first_sentence_ms and that sentence as first_sentence. If no space ever comes, the whole thing is one sentence.

If you cut by looking only at the period, you go wrong at '3.5' or 'a.m.'. The moment a space arrives after the punctuation is a safer signal that "the sentence has ended." Small models often violate the "three sentences" instruction — that is why you handle length in code.

Stop generation when someone breaks in

Stream 'List the numbers from 1 to 300, separated by commas.' with max_tokens=256, close the connection after 0.5 seconds, and write the number of content chunks received up to then as received_tokens and the ms until is_processing in /slots becomes false after closing as slot_idle_after_ms, to /root/voice/llm/cancel.json (max_tokens, received_tokens, closed_at_ms, slot_idle_after_ms).

Calling close() on the http.client connection is the "stop." If 'cancel task' appears in the server log (voice-llm log), the server understood. Choose a request that will surely run long — if generation finishes before you cut it, it is not a test.

Report

Copy into /root/voice/llm/report.json ttft_p50_ms and total_p50_ms (ttft.json), ttft_12docs_ms and prompt_n_12docs (the last entry of prefill.json), cache_ttft_ms (second in cache.json), first_sentence_ms (sentence.json), and cancel_idle_ms (slot_idle_after_ms in cancel.json).

Read the files from the previous steps and copy them. If you put the numbers side by side, you can see where to cut to open your mouth sooner.