Voice AI Agents — a pipeline that listens, looks things up and speaks
Is the caller done? Split turns with VAD and find barge-ins
Goal
Run an energy VAD and a neural VAD on the same call and count false alarms, write a rule that groups VAD segments into turns and test it on a call it has not seen, and find the time the person broke in from a microphone with echo mixed in.
Why it matters
Lateness in the endpoint decision is an answer latency that does not shrink even if you make the model fast. Set it short and you cut in the middle of speech; set it long and it is slow. The call in this lab (/opt/lab/fixtures/voice/vad/call_a.wav) was built by placing synthesized sentence pieces at fixed times, so the "actually spoken stretches" are in call_a.truth.json. The grader runs your turns.py again on a second call that you have not seen — code whose threshold was fitted to a single file falls over there.
Steps
- Write to
/root/voice/vad/energy.jsonthe result of an energy VAD that joins the stretches where the dBFS of 30 ms (480-sample) frames is greater than -35. - Write the segments of the silero VAD (
voicekit.models.vad(threshold=0.5, min_silence=0.1, min_speech=0.1), feeding 512 samples at a time) to/root/voice/vad/silero.json. - Overlay the two results on
speech_segmentsofcall_a.truth.json, countfalse_alarmsandmissed, and create/root/voice/vad/score.json. - Create
/root/voice/vad/turns.pycontainingdetect_turns(samples, sr, gap_s)— it joins the silero segments whose gap is shorter than gap_s and returns[(시작, 끝), …](the placeholders are the start and the end). - Create
/root/voice/vad/sweep.json, counting the turns of call_a at gap 0.2 · 0.3 · 0.5 · 0.8 · 1.2 seconds. - Using the shortest gap among those where the turn count is correct, write the answer latency for each turn — (the speech end seen by the VAD + gap) − the actual speech end — to
/root/voice/vad/latency.json. - From
/opt/lab/fixtures/voice/bargein/mic.wav, estimate the echo delay and scale ofplayback.wavand subtract them, save/root/voice/vad/residual.wav, and write the first speech time in the original and after subtraction to/root/voice/vad/bargein.json. - Create
/root/voice/vad/report.jsoncollecting the numbers from the previous steps.
Notes
- The silero VAD accepts a 512-sample (32 ms) window at 16 kHz. Feed it with
v.accept_waveform(x[i:i+512]), callv.flush()at the end, and then take the results out withwhile not v.empty(): v.front.start, v.front.samples; v.pop(). - A prediction counts as a hit if it overlaps the ground-truth segment "even a little." A predicted segment that does not overlap the ground truth at all is a false alarm, and a ground-truth segment that does not overlap any prediction is a miss.
- Common mistakes: memorizing gap for a single file (grading is done on another call), and looking for the echo delay in a stretch after the person has started speaking (the person's voice blurs the correlation).
- Documents: silero-vad · sherpa-onnx VAD · Stivers et al., 2009, Universals and cultural variation in turn-taking in conversation, PNAS
Energy VAD — every loud sound is speech
Split /opt/lab/fixtures/voice/vad/call_a.wav into 30 ms (480-sample) frames and write the stretches where consecutive frames have a dBFS greater than -35 to /root/voice/vad/energy.json as [{"start_s": …, "end_s": …}] (to two decimal places). Discard the leftover piece at the end.
The start of frame i is i×480/16000 seconds. When a speech frame begins, remember its number, and close the stretch at the start time of the frame where it ends. voicekit.dbfs gives you the RMS dBFS.
The segments of the silero VAD
Feed call_a into voicekit.models.vad(threshold=0.5, min_silence=0.1, min_speech=0.1) 512 samples at a time, call flush(), and write the resulting segments to /root/voice/vad/silero.json in the same shape (start = front.start ÷ 16000, end = start + len(front.samples) ÷ 16000).
Segments accumulate in a queue inside the VAD. Read front and pop() until empty() becomes true. If you do not call flush(), the segment still open at the end of the file does not come out.
Overlay the ground truth and count false alarms and misses
Taking speech_segments of call_a.truth.json as the ground truth, write the false_alarms (the number of predicted segments that do not overlap the ground truth at all) and missed (the number of ground-truth segments that do not overlap any prediction) for each of energy and silero to /root/voice/vad/score.json.
Two segments (a0,a1) and (b0,b1) overlap = a0 < b1 and b0 < a1. How did the energy VAD treat the door knock at 6 seconds?
Group VAD segments into turns
In /root/voice/vad/turns.py, create detect_turns(samples, sr=16000, gap_s=0.5). Go through the segments of the silero VAD (threshold 0.5 · min_silence 0.1 · min_speech 0.1) in order; if the gap from the end of the previous turn is shorter than gap_s, join it on, and otherwise start a new turn, and return [(시작 초, 끝 초), …] (the placeholders are the start and end in seconds). The grader runs both call_a and a call it has not seen with gap 0.5.
If you separate the work of cutting finely (VAD) from the work of grouping (gap), you can experiment by changing only gap. Put the executable code under if name == 'main': — the grader imports this file as a module.
Sweep gap — cutting and joining
With turns.py, count the turns of call_a at gap 0.2 · 0.3 · 0.5 · 0.8 · 1.2 seconds and write them to /root/voice/vad/sweep.json as {"0.2": {"turns": n}, …} (keys are strings).
There are 3 ground-truth turns in the turns of truth. If there are more turns than that, it cut at a pause in the middle of speech, and if fewer, it joined through to another turn.
The time waited before handing over the turn
Pick as gap_s the shortest gap among those in 05 whose turn count equals the ground truth, and create /root/voice/vad/latency.json with (VAD 가 본 턴 끝 + gap_s) − 실제 턴 끝 for each turn in per_turn_delay_s (the placeholders are the turn end seen by the VAD and the actual turn end), and their mean in mean_delay_s.
The moment the turn is judged to have ended is when it has stayed quiet for gap after the VAD saw the end of speech. The actual turn end is in the turns of truth.
Subtract the echo and find the time of the barge-in
Using /opt/lab/fixtures/voice/bargein/mic.wav (user + speaker echo) and playback.wav (the sound we sent out), find the delay with the largest cross-correlation (0–200 ms) in the first 1.5 seconds and the least-squares scale, and save /root/voice/vad/residual.wav with the echo subtracted. In /root/voice/vad/bargein.json, write echo_delay_ms, echo_gain, naive_first_speech_s (the time silero first picked up speech in the original microphone), and barge_in_s (the time it first picked up speech in the residual).
For each delay d, measure np.dot(mic[d:], play[:n-d]) and pick the largest d. The scale is g = (mic·ref)/(ref·ref). Use only the stretch before the user speaks, so that the person's voice does not blur the estimate.
Report
Copy into /root/voice/vad/report.json energy_false_alarms and silero_false_alarms (score.json), gap_s and mean_delay_s (latency.json), turns_call_a (the number of ground-truth turns), and barge_in_s (bargein.json).
Just read the files from the previous steps and copy them as they are. If you write them by hand, the rounding differs and they will not match.