Voice AI Agents — a pipeline that listens, looks things up and speaks
48 kHz source to 16 kHz frames — measure aliasing and stop it
Goal
Convert a 48 kHz stereo recording into the shape a speech model accepts (16 kHz · mono · 16-bit · 20 ms frames), and confirm with numbers the folding sound created by resampling without filtering, as well as dBFS and SNR.
Why it matters
The most common accident in a voice pipeline comes not from the model but from the shape of the sound. Even if the sample rate, channels, or byte order are off, no error is raised and only the recognition rate quietly collapses. The source (/opt/lab/fixtures/voice/audio/caller_48k.wav) has a 12 kHz pilot tone deliberately mixed in — if you do not filter when you downsample to 16 kHz, this tone folds into a sound you hear at 4 kHz. The grader re-reads the files you saved and directly measures the level of the 4 kHz component, the per-frame dBFS, the actual SNR, and the byte layout.
Steps
- Read the header of
/opt/lab/fixtures/voice/audio/caller_48k.wavand writesample_rate,channels,sample_width_bytes,frames, andduration_sto/root/voice/audio/info.json. - Make mono from the average of the two channels and save it at 48 kHz as is to
/root/voice/audio/mono48k.wav. - Without filtering, take every third sample (
x[::3]) and save it to/root/voice/audio/naive16k.wav(16000 Hz). - Apply a low-pass filter that passes only what is below 8 kHz, then take every third sample and save it to
/root/voice/audio/mono16k.wav(16000 Hz). - Split
mono16k.wavinto 20 ms (320-sample) frames and write the RMS dBFS of each frame to/root/voice/audio/levels.csv(index,start_s,dbfs). - Write the stretches where the level stays below -40 dBFS for 200 ms (10 frames) or more to
/root/voice/audio/silence.jsonas[{"start_s":…,"end_s":…}]. - Mix in seed-7 standard normal noise (
np.random.default_rng(7).standard_normal(n)) to reach an SNR of 10 dB, save it to/root/voice/audio/noisy10.wav, and write the value re-measured from the saved file to/root/voice/audio/snr.json. - Create
/root/voice/audio/stream.pcm, which ismono16k.wavwritten out as consecutive s16le 20 ms frames (640 bytes), and/root/voice/audio/stream.json, which records the format. Pad the missing part of the last frame with 0.
Notes
- Reading and writing: the standard library
waveandnumpyare enough. You can also use the helperfrom voicekit import read_wav, write_wav, to_int16, dbfsincluded in the image (it does the mono averaging and the 16-bit conversion). - Read 16-bit samples with
np.frombuffer(raw, dtype="<i2")and divide by 32768. Reshape stereo as(프레임, 채널)(the placeholders are the frame count and the channel count). - Common mistakes: computing the length as
frames / sample_rate / channels(framesis already the number of channel groups), only adding the two channels (it overflows), and matching SNR with the amplitude ratio (10 dB becomes 5 dB). - Concepts: Nyquist–Shannon sampling theorem · ITU-T G.711 · Python wave
Read the header first
Open /opt/lab/fixtures/voice/audio/caller_48k.wav with wave and write sample_rate, channels, sample_width_bytes, frames, and duration_s (seconds, to three decimal places) to /root/voice/audio/info.json.
wave.open(...).getframerate(), getnchannels(), getsampwidth(), and getnframes() give the four values. frames is the number of "samples grouped one per channel," so the length is frames ÷ sample rate.
Average the two channels into mono
Average the two channels of the source into mono and save it at 48000 Hz as is to /root/voice/audio/mono48k.wav (16-bit).
It is the mean along axis=1 of the (frames, channels) array. If you only add, loud parts exceed 1.0 and are clipped. When writing, multiply by 32768, round, and clip to the range -32768 to 32767.
Try decimating without filtering
From mono48k.wav, take every third sample (x[::3]) and save it to /root/voice/audio/naive16k.wav (16000 Hz mono). This is a deliberately wrong method.
48000 ÷ 3 = 16000, so the sample rate is right if you just decimate. But where did the original 12 kHz tone go? The grading message tells you the level of the 4 kHz component.
Filter, then downsample
Apply to mono48k.wav a low-pass filter whose cutoff frequency is lower than the new Nyquist limit (8 kHz), then take every third sample and save it to /root/voice/audio/mono16k.wav (16000 Hz mono). The 4 kHz folded component must be at least 30 dB smaller than when you just decimated.
A windowed sinc is the simplest low-pass filter: h = 2·fc·sinc(2·fc·n)·hamming, where fc is the cutoff frequency ÷ 48000. Divide by the sum of h to set the gain to 1, and apply it with np.convolve(x, h, mode='same'). The more taps you use, the steeper the cutoff.
The level of 20 ms frames (dBFS)
Split mono16k.wav into 320-sample (20 ms) frames (discard the leftover piece at the end) and write index,start_s,dbfs for each frame to /root/voice/audio/levels.csv. dBFS is 20·log10(RMS), and complete silence is -120.
RMS is the square root of the mean of the squares. Measure on the real numbers obtained by dividing the samples by 32768, so that 0 dBFS is the maximum. If you use 10·log10, the value is halved (it is the RMS of amplitude, not power, so use 20).
Find the silent stretches
From levels.csv, write the stretches where the level stays below -40 dBFS for 10 frames (200 ms) or more to /root/voice/audio/silence.json as [{"start_s": …, "end_s": …}]. The end is "the start of the first frame that crossed the threshold," and if it continues to the end of the file, it is the end of the last frame.
Remember the number of the frame that dropped below the threshold, and check the length at the moment it comes back above the threshold. If you append one large value to the end of the list, the same code closes the last stretch too.
Mix in noise at an SNR of 10 dB
Scale and add noise np.random.default_rng(7).standard_normal(n) of the same length as mono16k.wav so that the SNR is 10 dB (power ratio), save it to /root/voice/audio/noisy10.wav, and write the SNR measured by re-reading the saved file to measured_db in /root/voice/audio/snr.json (target_db is 10).
SNR = 10·log10(P_signal / P_noise), where P is the mean of the squares. The scale to multiply the noise by is sqrt(P_signal / 10^(10/10) / P_noise). Saving rounds to 16-bit, so it is accurate only if you re-read it and measure as noisy − clean.
Cut me into 640-byte frames with a WebSocket
Create /root/voice/audio/stream.pcm by writing the samples of mono16k.wav out consecutively in s16le (16-bit, smaller byte first) in 20 ms pieces (320 samples = 640 bytes) (pad the last frame with 0), and write format ("s16le"), sample_rate, channels, frame_ms, frame_bytes, and frames to /root/voice/audio/stream.json.
Write only the samples, with no WAV header. The .tobytes() of an np.int16 array follows the machine's byte order, so pin the dtype to '