Voice AI Agents — a pipeline that listens, looks things up and speaks
Speaking — time to first audio and producing 'speakable text'
In one line
TTS speed is measured by RTF (synthesis time ÷ audio duration). Even if RTF is below 1, synthesizing the whole answer at once delays the first sound. If you cut it into sentences (or clauses) and synthesize starting from the first piece, the first sound arrives several times sooner. And TTS reads what is written — '911' becomes 'nine hundred eleven,' '14:30' becomes 'fourteen thirty,' and the citation mark '[kb-hours]' becomes 'B hours.' There has to be a separate step that turns text written for the screen into text to be heard by the ear.
Why this was needed
LabHub's conversation practice spent 2.3 seconds of TTS per sentence (measured 2026-09-23, GPU server). Because it synthesized the whole answer after receiving all of it, if the model wrote five sentences you had to wait 12 seconds — as the comment in that file says, it becomes "a recital, not a conversation." The TTS in this course is much faster on CPU (below), but the principle is the same. The time to the first sound is proportional to the length of the first piece.
How it works
Piper and VITS. The TTS in this image is Piper's en_US ljspeech medium voice. It converts text to phonemes with espeak-ng, and then a VITS (Kim et al., 2021) neural network generates the waveform directly. The output is 22,050 Hz (model card). The training data, LJ Speech, is about 13,100 clips read by one person and is in the public domain. VITS draws a noise term on each synthesis, so even the same sentence produces a slightly different waveform each time — which is why the grader does not compare your WAV sample by sample.
RTF and the first sound. On a 2-core CPU Pod (nuc1), this model's RTF was about 0.04 — it makes 10 seconds' worth in 0.4 seconds. The int8 quantized version of the same model, even though its file is one third the size, was three times slower at RTF 0.12. Dynamic quantization costs something to go back and forth between integers and reals during computation, and with this CPU and this combination of operators that cost outweighed the gain. It is an example of why you must not choose by size alone. If you synthesize the whole thing, the first sound of a three-sentence answer comes out when the entire synthesis finishes, but if you cut it into sentences and synthesize only the first sentence first, you wait only that long.
Sentence splitting. If you cut at every period, it breaks on 'Dr. Rivera' and '$3.50'. The rule is this — a sentence ends only when a space and a capital letter (or the end of the text) follow the end punctuation, and abbreviations followed by a name or number, such as 'Dr.', 'Mr.', and 'No.', are not ends. 'a.m.' may be the end of a sentence ('…at 7:30 a.m. Results take…').
Overlapping with the LLM. If you cut pieces as text arrives from the LLM stream and pass them to TTS, the earlier sentence is synthesized while the LLM writes the next one. Small models often break "answer briefly," so sentences can be long. So even without a sentence end, you cut at a clause before a comma (five words or more) or at the twelfth word, so that code decides the length of the first piece. There is one thing to watch for. The CPU of this Pod is 2 cores, and the LLM server also uses 2 threads. If you run them overlapped, the two jobs share the same cores — the first sound comes sooner, but each one is slower than when running alone. The gain from overlapping is greatest when there are spare cores (measured in module 8).
Building the text to be read (text normalization). If you synthesize with this image's TTS and transcribe it again with ASR, you can see what goes wrong (measured): "Call 911" → "NINE HUNDRED ELEVEN," "at 14:30 [kb-hours]" → "FOURTEEN THIRTY B HOURS," "09:00" → "NINE ZERO ZERO." So before synthesis you delete citation marks, turn 24-hour times into a.m./p.m. (14:30 → two thirty p m), turn long numbers into single digits (911 → nine one one), and turn amounts into words ($80 → eighty dollars). You measure the effect with the TTS → ASR round-trip WER — using ASR as the listener in place of a human ear.
Transport format. The synthesized 22.05 kHz sound is converted into the shape the transport layer promised (16 kHz · s16le · 20 ms) and sent out. If you send 22050 as is as though it were 16000, it becomes a sound 1.38 times slower and lower.
What it looks like in the field
What is often missed in TTS first-sound latency is the preparation time of the first synthesis. After the model is loaded, the first call is slow because it prepares the computation graph. So a service synthesizes a short text once at startup (which is why this course's answer key calls 'warm up' first). It is the same idea as LabHub sending silence to the ASR to load it in advance.
Building the text to be read depends on language and region. '14:30' is read as 'two thirty p m' in American English and as 'half past two' on British broadcasts. Dates ('9/10') mean different things in different countries. So you gather the rules in one place and add one line to the round-trip WER test each time you see a new shape. This course's rules are the bare minimum for one American English voice.
What you will do in the next lab
You synthesize one sentence and measure RTF, and build sentence splitting that does not break on abbreviations or decimals. You compare the first sound of a three-sentence text synthesized whole versus by sentence, and cut an LLM stream into pieces and run it overlapped with TTS. You build a function that converts to the text to be read, measure its effect by round-trip WER, and convert the synthesized sound into 16 kHz 20 ms frames.