Voice AI Agents — a pipeline that listens, looks things up and speaks
Search with a misrecognised question, and speak only answers whose grounding checks out
Goal
Build paragraph search over phone queries transcribed by ASR, and add query rewriting, a refusal threshold, and citation checking in turn, to build a RAG in which ungrounded statements never leave the mouth.
Why it matters
Voice queries come with misrecognitions, filler words, and follow-ups. A small LLM gets both the format and the content wrong. So code must decide "what to look for" and "what is safe to say." In the query file (/opt/lab/fixtures/voice/rag/queries.jsonl), the asr field is what this image's ASR actually transcribed from synthesized speech, and relevant and facts are the ground truth attached by hand. The grader imports your search, rewriting, and verification functions and tests them on hidden cases, and checks again whether the final utterances are backed by evidence (it does not ask the LLM).
Steps
- Split
/opt/lab/fixtures/voice/kb/*.mdby blank lines and, for each paragraph excluding the title (#) line, write{"id": "kb-hours#1", "doc": "kb-hours", "text": …}to/root/voice/rag/chunks.jsonl. - Embed the paragraphs with
voicekit.models.Embedderand save them to/root/voice/rag/embed.npy((number of paragraphs, 384)). - Create
/root/voice/rag/search.pyin whichsearch(query, k=3)returns[(문단 id, 코사인 점수), …](the placeholders are the paragraph id and the cosine score). - Create
rewrite(asr_text, context=None), which removes filler words, normalizes number words, and resolves follow-ups, in/root/voice/rag/rewrite.py. - Write the top-1 and top-3 recall and the top-1 score per query, for the original and the rewritten queries, to
/root/voice/rag/eval.json. - Write the score threshold that separates answerable from unanswerable questions to
/root/voice/rag/threshold.json. - Ask the LLM with json_schema only the questions above the threshold, and write the answers and sources to
/root/voice/rag/answers.jsonl. - Create
verify(), which checks an answer, in/root/voice/rag/verify.py, and write the final utterances, with failed answers replaced by extraction, to/root/voice/rag/final.jsonl.
Notes
- Fields of the query file:
id,asr(the transcribed text),context(the id of the previous query a follow-up leans on),relevant(the ground-truth document),answerable, andfacts. - Recall is counted over the 10 answerable questions. If the document of the top-1 paragraph is one of the ground-truth documents, it is a top-1 hit.
- LLM call:
from voicekit.llm import chat→chat(messages, max_tokens=80, json_schema=schema)returns (text, first-chunk ms, total ms, timings).voice-llm upcomes first. - Common mistakes: including unanswerable questions in the recall denominator, asking the model even the questions below the threshold, and not binding the citation id with grammar so that it can cite a document that was not retrieved.
- Documents: all-MiniLM-L6-v2 model card · llama.cpp server — response_format · Cormack et al., 2009, Reciprocal Rank Fusion
Split into paragraphs
Read /opt/lab/fixtures/voice/kb/*.md in name order, split by blank lines (\n\n), drop pieces that are empty or start with #, and write one line per paragraph to /root/voice/rag/chunks.jsonl with text (whitespace folded into single spaces), id (document name#sequence number, starting from 1), and doc.
The document name is the file name with .md removed. " ".join(p.split()) folds line breaks and repeated whitespace into single spaces.
Embed the paragraphs
Embed the text of chunks.jsonl in order with voicekit.models.Embedder().embed([...]) and save it with np.save to /root/voice/rag/embed.npy ((number of paragraphs, 384), row order = chunks.jsonl order).
Embedder returns vectors normalized to length 1. If the row order is off, the ids in the search results point to the wrong paragraphs.
The cosine search function
In /root/voice/rag/search.py, create search(query, k=3). Load chunks.jsonl and embed.npy in advance, embed the query, and return k items of [(문단 id, 점수), …] in descending order of dot product (= cosine) (the placeholders are the paragraph id and the score). Create the model only once, the first time it is called.
One line, VEC @ q, gives the cosine with every paragraph. np.argsort(-s)[:k]. If you base file paths on os.path.dirname(file), it can be called from anywhere.
Clean up the query
In /root/voice/rag/rewrite.py, create rewrite(asr_text, context=None). Lowercase it and keep only words, remove filler words (um · uh · oh · er · ah · like · okay · ok · o · k · hi · hello · so · well · ohi), and turn number words into digits (eight → 8, twenty four → 24). If the result is 'and what about …', 'what about …', or 'how about …' and context exists, clean up context with the same rules, then replace the day with the new day if there is one, and otherwise append it at the end. The grader tests it with hidden cases.
The context of a follow-up is also an ASR result, so it goes through the same rewrite first. For numbers, add the tens (twenty…) and the ones (one…nine) when they are consecutive. 'o' and 'k' are what 'OK' was transcribed as 'O K'.
Measure recall and scores
Run search(…, 3) on the 12 queries with the original asr (raw) and after cleaning with rewrite (rewritten, with the asr of the previous query as context), then write recall_at_1 and recall_at_3 over the 10 answerable ones to /root/voice/rag/eval.json, and write the top-1 score of the rewritten queries in scores, by query id.
Convert paragraph ids to document names (doc) and compare with the ground-truth documents. q11 and q12 are unanswerable questions, so they are left out of recall, but write their scores — they are used for the threshold in the next step.
The threshold that separates unanswerable questions
Choose a threshold tau between the lowest score of the answerable questions and the highest score of the unanswerable questions in the scores of eval.json, and write it to /root/voice/rag/threshold.json (tau, answerable_min, unanswerable_max). All answerable questions must be at or above tau, and the unanswerable ones below tau.
The midpoint of the two values is a safe choice. If the two overlap, no threshold can get both right — in that case, look again at the query (the rewriting rules), not at the retrieval.
Ask the model only the questions above the threshold
For each query, rewrite and search (top-3); if the top-1 score is below tau, do not ask the model, and set answer to the refusal text ('I'm not sure about that. Let me connect you to the front desk.') and source to none. If it is above, put the retrieved paragraphs into the system prompt with [document id] attached, ask with a json_schema of answer (string) and source (an enum allowing only the retrieved document ids and none), and write the result. In /root/voice/rag/answers.jsonl, write one line each with id, query, retrieved (the retrieved document ids, without duplicates), score, answer, source, and gated.
schema = {"type": "object", "required": ["answer", "source"], "properties": {"answer": {"type": "string"}, "source": {"type": "string", "enum": retrieved + ["none"]}}}. What the grammar blocks is only the format — the content is checked in the next step.
Speak only what passes the check
In /root/voice/rag/verify.py, create verify(answer, source, retrieved, kb) — kb is {document id: full text of the document}, and it returns (whether it passed, the reason). Fail any answer that is empty, whose source is none, that cites a document that was not retrieved, or that has a number not in the document (including forms like the time 7:30). Then check answers.jsonl, and write to /root/voice/rag/final.jsonl (id, say, source, status): answers that passed as they are (status llm), questions refused at the threshold as the refusal text (refused), and failed answers changed to read out verbatim (extractive) the sentence in the cited document (or the top-1 document) whose embedding is closest to the question.
Extract numbers with re.findall(r"\d+(?::\d+)?", …) and check that all of them are in the document's list of numbers. Split sentences with (?<=[.!?])\s+. Extraction cannot be wrong, but it may not fit the question exactly — module 9 measures that cost.