LabHub
Get started
Learn Learning paths Courses

Voice AI Agents — a pipeline that listens, looks things up and speaks

From partial to final results — measuring streaming ASR and Whisper with WER

Continue in LabHub

Goal

Build a WER calculator yourself, compare streaming ASR and offline ASR on the same utterances, and measure accuracy and the latency "from end of speech to the final result." Confirm how the telephone band, a wrongly stated sample rate, and noise change WER.

Why it matters

The moment a voice assistant starts its answer is tied to "when the final result comes out." Streaming transcribes while the person speaks, so the final result comes out right away, while offline starts only when speech has ended. The accuracy number (WER) varies greatly with normalization, the way you aggregate, and the evaluation data, so you can trust the number only if you build the calculator yourself. The grader tests your wer.py on hidden cases, and runs the same model again to cross-check against your results.

Steps

  1. Create /root/voice/asr/wer.py in which wer(ref, hyp) returns {S, D, I, N, wer} (normalization: lowercase, strip punctuation, keep apostrophes).
  2. Feed /opt/lab/fixtures/voice/librispeech/ls01~ls08.wav into the streaming ASR 100 ms at a time and write the partial results, final result, and final latency to /root/voice/asr/stream.jsonl.
  3. Write the corpus WER of the final results to /root/voice/asr/wer_stream.json.
  4. Transcribe the same 8 with Whisper tiny.en and create /root/voice/asr/whisper.jsonl and /root/voice/asr/wer_whisper.json.
  5. Write the median "from end of speech to the final result" of the two methods to /root/voice/asr/latency.json.
  6. Compare the WER of the 8 kHz phone recording (phone_8k.wav) with the 16 kHz original (ls04), also measure it when you falsely say 8 kHz is 16000, and write it to /root/voice/asr/phone.json.
  7. Write the corpus WER when white noise at SNR 20 · 10 · 0 dB is mixed in to /root/voice/asr/noise.json.
  8. Collect the measurements and write which setup you would choose to /root/voice/asr/report.json.

Notes

Build a WER calculator

In /root/voice/asr/wer.py, create wer(ref, hyp). Lowercase both texts, remove punctuation other than apostrophes, compute the word-level minimum edit distance, and return {"S": 치환, "D": 삭제, "I": 삽입, "N": 정답 단어 수, "wer": (S+D+I)/N} (the placeholders are the substitution count, the deletion count, the insertion count, and the number of reference words; if the reference is empty, wer is 0). The grader tests it with 8 hidden cases.

Do Levenshtein distance with words instead of characters. After filling the table d[i][j], trace back from the bottom right and count which edits they were. The apostrophe in "don't" is part of the word, so it must be kept.

Feed it 100 ms at a time and record partial results

Feed ls01 to ls08 from /opt/lab/fixtures/voice/librispeech/refs.tsv into the streaming ASR 1600 samples (100 ms) at a time, collect {"audio_s": 지금까지 넣은 소리 길이, "text": 부분 결과} (the placeholders are the length of audio fed so far and the partial result) every time the result changes, measure the ms from the moment you feed the last sound until the final result comes out as final_ms, and write one line per utterance (id, partials, final, final_ms) to /root/voice/asr/stream.jsonl.

You must append 0.66 seconds of silence after the last chunk and call input_finished() for the final word to come out. final_ms is the time that finish took — the share the user waits after speech has ended.

Corpus WER

With your wer.py, add up the error counts and word counts of all 8 utterances to get the corpus WER and write it to /root/voice/asr/wer_stream.json as S, D, I, N, and wer.

It is not the average of per-file WER. Compute it as the sum of errors ÷ the sum of words, so that longer utterances are weighted more heavily.

Compare with Whisper tiny.en

Transcribe the same 8 utterances one at a time with voicekit.models.whisper(), create /root/voice/asr/whisper.jsonl (id, text, compute_ms, audio_s), and write the corpus WER to /root/voice/asr/wer_whisper.json.

An offline recognizer is s = rec.create_stream(); s.accept_waveform(sr, x); rec.decode_stream(s); s.result.text. Whisper adds punctuation and capitalization, so your normalization does its work here.

From end of speech to the final result

Write the median (nearest-rank: the ⌈0.5·n⌉-th value after sorting) of final_ms in stream.jsonl and of compute_ms in whisper.jsonl to /root/voice/asr/latency.json as stream_final_ms_p50 and whisper_ms_p50.

Whisper starts only after speech has ended, so the whole compute_ms is the user's wait. Streaming finished most of its work while the person was speaking, so only final_ms is waited.

The 8 kHz telephone band — and what if you lie about the sample rate

Transcribe /opt/lab/fixtures/voice/librispeech/phone_8k.wav (8 kHz), which has the same sentence as ls04 (16 kHz), with the streaming ASR, and write the WER against the ls04 reference as wer_16k and wer_8k, and the WER of the result of feeding the 8 kHz samples with accept_waveform(16000, …) as wer_8k_as_16k, to /root/voice/asr/phone.json.

The first argument of accept_waveform is the sample rate of the sound. If you say 8000, sherpa-onnx converts it to 16 kHz internally. If you falsely say 16000, 8,000 samples are read as 0.5 seconds and the sound runs twice as fast.

The stronger the noise

Mix noise np.random.default_rng(100 + k).standard_normal(n) into ls01 to ls08 (k is the file's index in sorted order) at SNR 20 · 10 · 0 dB (power ratio), transcribe it with the streaming ASR (100 ms chunks), and write the corpus WER to /root/voice/asr/noise.json as {"20": …, "10": …, "0": …}.

The scale is the same as in module 1: sqrt(P_signal / 10^(SNR/10) / P_noise). If you fix the seed per file, the grader can regenerate the same noise and cross-check.

What to choose

In /root/voice/asr/report.json, copy wer_stream, wer_whisper, stream_final_ms_p50, whisper_ms_p50, wer_8k, and wer_0db from the files of the previous steps, and write one of stream, whisper, and two-pass in choice.

Look at latency, accuracy, and similarity of the data together. Do not forget that this material (LibriSpeech) is the same as the streaming model's training data.