LabHub
Get started
Learn Learning paths Courses

Voice AI Agents — a pipeline that listens, looks things up and speaks

Sound as numbers — sample rate, frames, dBFS and sound that folds back in

Continue in LabHub

In one line

What a microphone gives you is a sequence of air-pressure numbers (PCM) measured at regular intervals. How many times it measured per second is the sample rate, and the size of one number is the bit depth. Every part of a voice pipeline passes this number sequence along in frames of around 20 ms. If you do not filter when you change the sample rate, sounds that were never there appear, and if you measure level in dBFS, a single threshold can separate silence from noise.

Why this was needed

When you first build a voice assistant, you choose the model first. But the first accident in the field comes not from the model but from the shape of the sound. A browser records at 48 kHz, the telephone network sends at 8 kHz, this course's recognition model takes 16 kHz, and the synthesis model outputs 22.05 kHz. All four numbers show up within a single call. If you write the sample rate wrong in one place, the sound runs twice as fast and recognition drops to zero (module 3 measures this for real — pretend 8 kHz is 16 kHz and WER is 100%), and if you get the byte order wrong, only noise is left. No error is raised. The numbers are still numbers.

How it works

Sample rate and the Nyquist limit. If you measure fs times per second, you can represent only components up to fs/2 Hz (the sampling theorem). Measured at 16 kHz, that is up to 8 kHz. Most of the components of human speech that matter for understanding lie below that, so speech recognition models use 16 kHz as the standard. Landline telephony (ITU-T G.711) measures at 8 kHz and carries only 300–3,400 Hz — which is why telephone voices sound muffled.

Folding (aliasing). When you go from 48 kHz down to 16 kHz by just taking every third sample, components that were above 8 kHz do not disappear but fold down into the lower range. A 12 kHz component shows up at 16 − 12 = 4 kHz. It is a beep that was not in the original. That is why, before decimating, you first apply a low-pass filter that cuts below the new Nyquist limit (8 kHz). The lab source has a 12 kHz pilot tone (-26 dBFS) mixed in so that you can see this. If you just decimate, -29 dBFS appears at 4 kHz, and if you filter with a 127-tap windowed sinc filter before decimating, it drops to -65 dBFS (measured on this image). Conversely, raising 8 kHz to 16 kHz does not create sound above 4 kHz — information that was not there is not created.

Bit depth and byte order. 16-bit PCM is an integer from -32768 to 32767. Calculations are done on real numbers in [-1, 1) obtained by dividing by 32768, and integers are what you write to files and the network. Bytes are usually ordered with the smaller one first (little-endian), so the format name is s16le. When you mix two channels down to mono, you take the average. If you only add them, two channels at 0.8 become 1.6 and are clipped.

Frames and level. Sound is handled in cuts of 20 ms (320 samples at 16 kHz, 640 bytes). The properties of human speech stay almost the same over tens of ms, and real-time transport (such as Opus in WebRTC) also commonly uses 20 ms. The level of a frame is written as the RMS (root mean square) in dBFS — a logarithmic scale with the maximum value (1.0) set at 0 dB: 20·log10(RMS). -40 dBFS is 1% of the maximum, and -60 dBFS is 0.1%. Because multiplication becomes addition, a single threshold can separate things like "speech is near -20, room sound is near -50."

SNR. The signal-to-noise ratio is a ratio of power: 10·log10(P_signal / P_noise). Written as an amplitude ratio, it becomes 20·log10. If you mix the two up, you aim for 10 dB and end up with 5 dB or 20 dB — that is exactly the counterexample in the lab.

What it looks like in the field

This is the boundary with this course's real-time communication course. That course carries the browser microphone over a WebSocket to the server, and this course starts from the promise that "16 kHz · mono · s16le · 20 ms frames" arrive. If any one of the four slots in that promise is off, the receiver hears the wrong sound without any error. So even if you do not carry the format in every message, it is good to exchange it once when the connection opens so that both sides confirm it.

LabHub's language conversation practice feature takes the opposite approach — it records a whole sentence with the browser MediaRecorder and sends it (backend/app/lang_talk.py, static/lang/talk.js). The upper limit on a single recording was set at "2 MB, with headroom over 15 seconds × 16 kHz × 16 bits," and that came from this calculation. Sending it whole keeps the implementation simple, but recognition starts only after the person has finished speaking. The reason to send in frames is to remove that waiting time (modules 3 and 8).

Frame size has a trade-off too. The smaller the frame, the less time it takes to collect one piece (20 ms for a 20 ms frame), but the number of messages grows and the share taken by headers and system calls gets bigger. Conversely, if you raise it to 100 ms, messages drop by a factor of five, but every stage receives sound at least 100 ms later. This course's streaming ASR takes 100 ms chunks — effectively collecting five WebSocket frames and feeding them in at once.

What you will do in the next lab

You read the header of a 48 kHz stereo source, mix it down to mono, downsample to 16 kHz without filtering and measure the 4 kHz folding, then downsample properly with a low-pass filter. You find the silent stretches using the dBFS of 20 ms frames, mix in fixed-seed noise at an SNR of 10 dB, and then cut it into 640-byte frames through a WebSocket and save them.