Voice AI Agents — a pipeline that listens, looks things up and speaks
When to answer — VAD, end-of-turn detection, barge-in
In one line
When a voice assistant opens its mouth is decided by the judgment that speech has ended (endpointing). VAD (voice activity detection) separates, frame by frame, whether a sound is speech or not, and turn detection groups the VAD segments into turns by "how long has it been since the speech stopped." If you set the pause threshold (gap) short, you cut in the middle of speech, and if you set it long, you answer late. If a person breaks in while the assistant is speaking (barge-in), you must first subtract the echo of the speaker sound coming back into the microphone, so that the person's voice becomes visible.
Why this was needed
In conversation between people, the gap at which the turn passes is short. A study that measured conversations in 10 languages (Stivers et al., 2009, PNAS) reported that the gap before an answer begins clusters around 0–200 ms. People predict the end before the other person has finished speaking. Machines cannot, and only after confirming it has gone quiet do they know it has ended. That confirmation time is the minimum latency of the answer — it does not shrink however fast you make the model (module 8 measures this share separately).
A single threshold does not solve it either. In "I booked it for Tuesday... (a thinking pause) ...I want to change it to Thursday," the pause can exceed 0.5 seconds. If you cut there, the assistant ends up answering "I booked it for Tuesday."
How it works
Energy VAD. If the dBFS of a frame is above a threshold (for example -35), it counts as speech. It is fast and easy to explain, but every loud sound is speech. The lab's call has a short noise of someone knocking on a door at the 6-second mark, and the energy VAD picks this up as one speech segment (a false alarm). Conversely, it misses quiet speech.
Neural VAD (silero). A small neural network outputs a "probability of speech" for each frame (512 samples at 16 kHz, 32 ms); above the threshold it counts as speech, and when it is quiet for at least min_silence_duration it closes the segment. The silero VAD in this image (onnx, MIT) processed 16.7 seconds of sound in 56 ms on a 2-core CPU Pod (measured on nuc1). In the lab material it did not pick up the door knock as speech. However, it closes a segment about 0.5 seconds later than the real end of speech (measured at 0.5–0.7 seconds in this material) — because it counts the tail of the sentence and the breathing as speech. This lateness is added directly to the answer latency as well.
Grouping into turns. The VAD cuts a segment at every pause. A turn is those segments joined together wherever the gap is shorter than gap. If you set gap to 0.2 seconds, it cuts turns even at a thinking pause (cutting in the middle of speech), and if you set it to 1.2 seconds, it either joins through to the next person's turn as one turn or waits over a second every time. In the lab you sweep gap from 0.2 to 1.2 seconds, find the range where the turn count is correct, and measure the answer latency at the shortest value in it — (the speech end seen by the VAD + gap) − the actual speech end.
In practice you add a semantic signal on top. You look at whether the ASR's partial result ends like a sentence ("...I want to change it"), or whether the intonation has fallen, and shorten or lengthen gap accordingly. Streaming ASR models have endpoint rules too — rule2 in sherpa-onnx is "after something has been recognized, it ends if quiet for N seconds."
Barge-in and echo. While the assistant is speaking, the microphone picks up not only the person's voice but also the assistant's own voice leaking in from the speaker. If you run the VAD in this state, it judges "the person has broken in" the moment playback starts and cuts itself off. The solution is to use the fact that we know the sound we sent out (the reference signal). Treat the microphone as person + g·(the reference delayed by d), find the d with the largest cross-correlation and the least-squares scale g in the stretch before the person speaks, and subtract; the person's voice is what remains. A real echo canceller (AEC) is an adaptive filter that follows even the room's reflections, and in a browser you turn it on with the echoCancellation constraint of getUserMedia. The echo in the lab is a simple model made with one delay and one scale, so you see only the principle.
What it looks like in the field
LabHub's conversation practice leaves the endpoint decision to the person — you press and release the record button (MediaRecorder in static/lang/talk.js). It cannot go wrong and the implementation is simple, but the person finishes speaking, presses the button, and only then is the whole recording sent. A hands-free voice assistant like a phone line has to make this decision by machine, and at that moment this module's trade-off arises.
The most common mistake in choosing a threshold is memorizing it for a single recording. Even looking at just the two calls in this course, in one call the gaps between VAD segments separate comfortably into 0.29 seconds (mid-speech) and 1.06 seconds (turn change), but in the other call the gap at a turn change narrows to as little as 0.73 seconds. If you pick 0.74 seconds from one file, the two people's turns get joined into one turn in the other file. So you test the rule on material it has not seen, and in production you collect the handed-over turns and the cut-off turns as samples and revisit the threshold.
What you will do in the next lab
On call call_a, you overlay the segments of the energy VAD and the silero VAD on your own ground truth and count false alarms and misses. You write turns.py, which groups VAD segments into turns, and the grader tests whether it also works on a call it has not seen, and you sweep gap to measure the range where turns are correct and the answer latency. Finally, from a microphone with echo mixed in, you estimate the delay and scale, subtract them, and find the time the person broke in.