LabHub
Get started
Learn Learning paths Courses

Voice AI Agents — a pipeline that listens, looks things up and speaks

Transcription — streaming transducers, Whisper, and what WER hides

Continue in LabHub

In one line

Streaming ASR keeps revising partial results while sound comes in, and when speech ends it produces a final result after a short finish. Offline ASR (Whisper) cannot start until it receives the whole utterance. Accuracy is measured by WER — (substitutions + deletions + insertions) ÷ number of reference words — but this number swings a great deal with how closely the evaluation data resembles the model's training data.

Why this was needed

A voice assistant's answer begins after the judgment that "speech has ended" (module 2). If the recognition result is already almost complete at that moment, you move straight on to the next stage, and if recognition starts only then, you wait in proportion to the length of the utterance. That difference is the reason to use streaming. At the same time, a partial result serves as a signal that "it is listening" and as grounds for starting a search in advance.

But WER, the first number you grab when choosing a model, hides more than you would expect. There was a real experience while choosing the model to put in this image. The 20M streaming model chosen first was fine on public evaluations, but when fed this course's queries it lost the first 1–2 seconds of the stream entirely — "How late can I cancel without paying the fee?" came out as "LE WITH THAT PAIN THE FEET." When the same file was concatenated twice, the second copy was accurate. It was a failure invisible in the average WER, and it is why the model was changed (recorded in labs/voice/Dockerfile).

How it works

Transducer (RNN-T). The streaming model in this image is an icefall zipformer transducer (trained on LibriSpeech, Apache-2.0). The encoder takes a chunk of sound and produces a representation, and the joiner decides "what to emit at this point (or nothing at all)." The encoder was trained to see only chunks of a fixed size and a limited left context (chunk-16 · left-128 in the model name), so it produces results chunk by chunk without waiting for the future. That is why partial results appear, and for the last chunk a short stretch of silence (0.66 seconds in this course) is appended to make it finish. This finish takes tens of ms — in this image, the median "from the last sound to the final result" over 8 utterances was 44 ms (Docker on a Mac, 2 CPUs).

Whisper. OpenAI's Whisper (Radford et al., 2022) is an encoder-decoder model trained on 680,000 hours of web data. It converts sound into a log-mel spectrogram of a 30-second window and feeds it to the encoder, and the decoder writes the text one token at a time. It writes after seeing the whole utterance, so the result comes with punctuation and capitalization, but it cannot start until speech has ended. Under the same conditions, one utterance took 734 ms (median). The int8 version of the smallest one, tiny.en (39 million parameters), was used.

Partial results fluctuate. A streaming model may revise earlier words after seeing later context. If you show partial results on screen, the user sees the text change. So it is common to show only the "stable leading part" in bold, or to do hard-to-undo things like search only on the final result.

WER and normalization. WER is the word-level edit distance. Always normalize before comparing — case, punctuation, apostrophes. If you compare "Sunday?" and "SUNDAY" without normalization, you end up grading the notation rules, not the model. Corpus WER is not the average of per-file WER but the sum of errors ÷ the sum of words. It is meant to keep the 100% error of one short utterance from ruining the average.

WER depends on the data. Putting three numbers measured in this image side by side makes it clear. On 8 LibriSpeech test-clean utterances, the WER of the streaming zipformer was 0.95% and that of Whisper tiny.en was 6.7% — because the zipformer was trained on that very corpus. But on 12 phone queries made with synthesized speech (module 9), the same zipformer got almost 20% wrong ("GUYS" → "GUISE", "refill" → "ORIFYL"). The more the evaluation data resembles the training data, the better the number looks. When choosing a model, you have to measure with data that resembles your service.

Be honest about the sample rate. sherpa-onnx converts on its own when it receives sound that is not 16 kHz. If you tell it an 8 kHz phone recording is 8000, the WER was 0% on this lab's material, but if you falsely say 16000, the sound runs twice as fast and 100% is wrong. Noise is different. When white noise was mixed into the same 8 utterances, the WER rose to 0% at an SNR of 20 dB, 2.9% at 10 dB, and 24.8% at 0 dB.

What it looks like in the field

Two-pass recognition. A common setup is to show partial results immediately with a streaming model, and when speech ends, transcribe once more with a larger offline model to correct the final result. You spend on the second pass only as much as the latency budget allows. The ASR in LabHub's conversation practice calls a large model loaded on the GPU once at the end of the recording, and the first request took 78.7 seconds when the weights were offloaded and 0.4 seconds when they were loaded — so at the moment the user takes their seat, it sends silence to load it in advance (warm_asr in backend/app/lang_talk.py). When you use a large model, "warm-up" becomes part of the design.

What you will do in the next lab

You build a WER calculator and have it tested against hidden cases, receive 8 LibriSpeech utterances in 100 ms streaming chunks, and record the partial results and final latency. You transcribe the same utterances with Whisper and compare WER and latency, measure WER for the 8 kHz telephone band, for a falsely declared sample rate, and for different noise levels, and then write in a report which setup you would choose.