Voice AI Agents — a pipeline that listens, looks things up and speaks
Search by voice — misrecognition, asking back, refusing, and wrong answers that pass the grounding check
In one line
Voice queries are short, conversational, and arrive with the ASR's mistakes still in them. So before searching you clean up the query (strip filler words, normalize number words, resolve follow-ups that lean on the previous question), refuse unanswerable questions first using the retrieval score, and speak the model's answer only after checking it against the document it cited. The check can only see whether the numbers are in the document — you have to evaluate knowing that wrong answers that pass the grounding check remain (module 9).
Why this was needed
The RAG in the LLM engineering course takes questions typed as text. A phone is different. The 12 queries in this module are what this image's streaming ASR actually transcribed from synthesized speech (fixtures/voice/rag/queries.jsonl). "What time do you guys open on Saturday?" came out as "OHI WHAT TIME DO YOU GUISE OPEN ON SATURDAY," and "How long does a refill take?" as "HOW LONG DOES ORIFYL TAKE." And the second question is "O K AND WHAT ABOUT SUNDAY" — without the previous question you cannot tell what is being asked.
How it works
Paragraph-level embedding. An answer to be read aloud over the phone is one or two sentences, so you split documents into paragraphs and embed each paragraph. The embedding model in this image is all-MiniLM-L6-v2 (qint8 ONNX). As the model card says, if you mask the token vectors with the attention mask, average them (mean pooling), and normalize to length 1, the dot product is the cosine similarity.
Query rewriting. Three rules cover most of this material. (1) Remove a list of filler words (um · uh · oh · like · okay · o k …). The list also includes 'ohi', which this ASR actually produced — collecting observed misrecognitions into a dictionary is the way it is done in the field. (2) Turn number words into digits (twenty four → 24). The documents write '24 hours'. (3) For a follow-up like 'and what about Sunday', take the previous question and swap only the day. Misrecognitions like 'ORIFYL' cannot be fixed by rules — you can only hope that the embedding finds the paragraph that is close in meaning.
When you measure. On these 14 documents (41 paragraphs), top-1 recall was 10 out of 10 even before rewriting. The embedding placed the question containing 'ORIFYL' close to the prescription refill paragraph. But when switched to a hybrid mixed with BM25 (RRF, k=60), top-1 dropped to 8 of 10 with the original queries and 7 with the rewritten queries — misrecognized words like 'VITIO' and 'ORIFYL' are in no document and so give BM25 no signal, and the remaining common words (how, long, the) blurred the ranking (values measured in this image). Hybrid does not always win. Where rewriting helped a lot was not retrieval but the meaning of the question. If you give "and what about sunday" to the model as is, it cannot tell what was asked.
Refusal threshold. Even for a question with no answer (a pizza place recommendation, the wifi password), retrieval still brings back something. If the score of the closest paragraph is lower than the threshold τ, you refuse without even asking the model — so as not to give a small model the chance to make things up. In this material the lowest score for an answerable question is 0.345 and the highest score for an unanswerable one is 0.260, so a threshold sits between them. The reason to be careful with rewriting rules showed up here too. At one point, when a 'you → the clinic' substitution was added, "can you recommend a good pizza place" became "can the clinic recommend…" and its score rose to 0.397, so no threshold could separate it from the answerable questions. So that rule was removed.
Format by grammar, content by checking. If you tell a 0.5B model to 'write the source like [kb-hours]', it breaks the format first. With llama-server's response_format (json_schema), you bind the output to {"answer": …, "source": …} and make source an enum that allows only the document ids found this time and none, so the model cannot cite a document that was not retrieved. Then code checks: are the numbers (times, amounts, days) in the answer present as they are in the cited document? In fact, this model produced "refill takes 10 business days" while citing the blood test document — the document says 3 business days. An answer that fails the check is not spoken, and the sentence in the cited document closest to the question is read out verbatim (extraction). It cannot be wrong, in exchange for being less natural.
Wrong answers that pass the grounding check. The number check lets through even an answer of weekday hours (8:30 a.m. to 6 p.m.) to "what time do you open on Saturday," because the numbers are in the document. In fact, this pipeline answered the Saturday question and the Sunday follow-up with weekday hours. Wrong answers like this cannot be stopped by a verifier and show up only when evaluated against an answer key (factual accuracy in module 9).
What it looks like in the field
Citations in a voice assistant are different from footnotes on a screen. You must not read '[kb-hours]' aloud (it is removed in module 7's step of building the text to be read). Instead, citations are kept in logs and evaluation, so that you can later ask "what did it base that on."
What you will do in the next lab
You split the KB documents into paragraphs, embed them, and build a cosine search function. You have your query rewriting rules tested against hidden cases, measure recall and scores with the original and rewritten queries, and then set the threshold that separates unanswerable questions. Only questions above the threshold are sent to the LLM as JSON, the verifier filters the results, and answers that do not pass are replaced with extraction to produce the final utterance.