LabHub
Get started
Learn Learning paths Courses

Voice AI Agents — a pipeline that listens, looks things up and speaks

Pull the first sound forward — sentence splitting, overlapped synthesis, speakable text

Continue in LabHub

Goal

Measure TTS RTF and first-sound latency, cut by sentence or clause and synthesize overlapped with an LLM stream, and convert text written for the screen into text to be heard by the ear, checking by round-trip WER whether it becomes easier to understand.

Why it matters

What the user waits for is not the whole answer but the first sound. The first sound is determined by the length of the first piece and the synthesis speed. And TTS reads what is written, so a '14:30 [kb-hours]' that looks fine on screen sounds odd to the ear. The grader does not compare the synthesized WAV sample by sample (VITS produces a slightly different waveform every time). It looks at consistency of length, RTF, and order of events, and checks whether it is intelligible by transcribing your WAV again with ASR.

Steps

  1. Synthesize one sentence, save it as /root/voice/tts/hello.wav, and write the synthesis time, length, and RTF to /root/voice/tts/synth.json.
  2. Create sentences(text), which does not break on abbreviations and decimals, in /root/voice/tts/split.py.
  3. Write the time to synthesize a three-sentence text whole and the time to synthesize only the first sentence to /root/voice/tts/first_audio.json.
  4. Cut an LLM stream into pieces and synthesize them overlapped with TTS, saving the event times to /root/voice/tts/stream.json and the concatenated sound to /root/voice/tts/stream_out.wav.
  5. Create speakable(text), which turns citation marks, 24-hour times, long numbers, amounts, and abbreviations into text to be read, in /root/voice/tts/normalize.py.
  6. Synthesize three sentences both as the original text and as text to be read, transcribe each with ASR, and write the result to /root/voice/tts/roundtrip.json.
  7. Convert stream_out.wav into 16 kHz s16le 20 ms frames and create /root/voice/tts/tts_stream.pcm and /root/voice/tts/tts_stream.json.
  8. Create /root/voice/tts/report.json collecting the measurements.

Notes

Synthesize one sentence and measure RTF

Create the model, warm it up once with a short text, synthesize 'Your appointment is on Tuesday at ten in the morning.', save it to /root/voice/tts/hello.wav (22050 Hz mono), and write sample_rate, synth_ms (the time taken by generate), audio_s, and rtf (synth_ms/1000/audio_s) to /root/voice/tts/synth.json.

Wrap only the generate call with time.perf_counter(). If RTF is below 1, it produces faster than the speaking speed. If you include model loading (1–2 seconds), RTF looks like it exceeds 1.

Sentence splitting that does not break

In /root/voice/tts/split.py, create sentences(text). After folding whitespace into single spaces, cut only when a space and a capital letter (or the end of the text) follow .!?, but do not cut after 'Dr.', 'Mr.', 'Mrs.', 'Ms.', 'St.', or 'No.'. Joining the pieces must give back the original, and an empty text gives an empty list. The grader tests it with hidden cases.

Find candidates with re.finditer(r".!?", text), and skip a candidate if the word just before it is in the abbreviation list. '$3.50' is not a candidate because a digit follows the period.

Whole versus by sentence — up to the first sound

Write to /root/voice/tts/first_audio.json the time to synthesize 'Please arrive fifteen minutes before your visit. Bring a photo ID and your insurance card. If you need to cancel, call at least one day before.' whole (whole_ms), the time to synthesize only the first sentence split by split.py (chunked_ms), and the number of sentences (sentences).

If whole, everything has to be made before the first sound can go out, and if by sentence, you only need to make the first sentence. Measure both with a warmed-up model.

Synthesize overlapped with the LLM

After voice-llm up, receive 'You are a clinic phone assistant. Answer in two or three short sentences.' and 'How do I prepare for a fasting blood test?' with voicekit.llm.chat_stream(msgs, max_tokens=80), and each time text accumulates, cut the leading piece (a sentence if one has ended, otherwise a clause before a comma of more than five words, otherwise twelve words) and pass it to the TTS thread through a queue. Write the events llm_first_token, chunk (the piece passed, including text), audio (a piece whose synthesis has finished, including audio_s), and llm_done on a single clock to /root/voice/tts/stream.json, and save the sound concatenated in order to /root/voice/tts/stream_out.wav (22050 Hz).

Run TTS separately in a threading.Thread and receive pieces through a queue.Queue. Put None at the end to finish the thread and join it. There must be two or more pieces for overlap to occur — even if the small model answers in one sentence, you can cut by commas or word count.

From text for the screen to text for the ear

In /root/voice/tts/normalize.py, create speakable(text). Delete citation marks [kb-…], turn 'H:MM a.m./p.m.' and 24-hour 'HH:MM' into words (14:30 → two thirty p m, 09:00 → nine o'clock a m), turn numbers of three or more digits into single digits (911 → nine one one), '$80' into eighty dollars, the remaining one- or two-digit numbers into words, 'Dr.' into Doctor, and 'a.m.' and 'p.m.' into 'a m' and 'p m'. The grader tests it with hidden cases.

The order matters. Convert the time (9:30) first so that the number rule does not read 9 and 30 separately. One small function that turns 0–99 into words serves times, amounts, and numbers alike.

Measure intelligibility by TTS → ASR round trip

Synthesize 'Your visit is at 14:30 [kb-hours].' · 'Call 911 right away.' · 'A visit costs $80.' each as the original text (raw) and converted with speakable (norm), convert to 16 kHz and prepend 0.3 seconds of silence, save them as /root/voice/tts/rt{i}_{kind}.wav (i is 1–3), and feed them into the streaming ASR 100 ms at a time to transcribe. Write six lines {"id", "kind", "tts_input", "heard", "meant"} (meant is speakable(original text)) to /root/voice/tts/roundtrip.json.

You use ASR as the listener in place of a human ear. The grader transcribes your WAVs again and compares the WER against meant for raw and norm. If the piece size differs, streaming results can vary slightly, so fix it at 100 ms.

Convert to the transport format

Convert stream_out.wav (22050 Hz) to 16 kHz with voicekit.resample_linear and create /root/voice/tts/tts_stream.pcm, joined as s16le 20 ms (320-sample) frames (pad the last frame with 0), and /root/voice/tts/tts_stream.json, which records format, sample_rate, channels, frame_ms, frame_bytes, and frames.

It has the same shape as the last step of module 1. If you send 22050 as is while calling it 16000, it becomes a sound 1.38 times slower and lower.

Report

In /root/voice/tts/report.json, write rtf (synth.json), whole_ms and chunked_ms (first_audio.json), and llm_first_token_ms, first_audio_ms (the first audio), and llm_done_ms, which are the first times of each event in stream.json.

Read the event list from the front and keep only the first time you see for each name (dict.setdefault).