Voice AI Agents — a pipeline that listens, looks things up and speaks
Tools driven by voice — state machines, confirmation questions, retries and handing off to a human
In one line
Accidents in a voice agent usually happen in the tools. A reservation gets created without a "yes" answer, a timeout gets resent and there are two reservations, or it repeats the same question without understanding. So what can be done is decided by a state machine, and the model does only the job of turning the user's words into slots (date, time, name). Irreversible tools are called only after a confirmation question, and timeouts are not resent. If it fails to understand twice, it hands off to a person, and when it does, it summarizes what it has already heard.
Why this was needed
If you build a phone reservation assistant with a single LLM, you give the model the whole conversation and also leave "what to do next" to it. When I had the 0.5B model in this image classify intents with json_schema, it kept the format perfectly but got only 1 of 8 sentences right — it classified almost every sentence as 'book,' the first value of the enum (you measure this yourself in step 2 of this module's lab). Enforcing the format with grammar does not enforce the content. If you leave "is it OK to book now?" to a model like this, a reservation can be made even on a "no."
How it works
Slot filling and state. A reservation needs three slots: date, time, and name. The state flows LISTEN → ASK_DAY → ASK_TIME → ASK_NAME → CONFIRM → DONE, and from any state it can go to HANDOFF (to a person). If the user says "Tuesday at 10 a.m., the name is Jamie" all at once, there are no empty slots, so it goes straight to CONFIRM. Because the state decides what can be done, book is called only when a "yes" is heard in CONFIRM — this holds structurally, whatever the model says.
A "no" to the confirmation question. "No, make it two thirty" is not a rejection but a correction. If it contains a new time, you change only that slot and confirm again, and if there is no information at all, you do not know what was wrong, so you ask for the time again from the start. Voice has more misrecognition than text, so reading back right before an irreversible action is not optional.
Retries and backoff. Tool errors are divided into two kinds. A timeout (ToolTimeout) may be fixed by sending the same request again. A rejection (ToolError, for example, that slot was just filled) gives the same answer even if you resend it. For a timeout on a read-only tool (finding open times), you call again up to two times, pausing for longer each time, such as 0.5 seconds and then 1.0 seconds, and if it still fails, you do not hold on but hand off. On the phone, a few seconds of waiting is silence.
Irreversible tools. If a reservation (book) ends in a timeout, that is not "it did not work" but "we do not know." The server may have created the reservation and only the response was late. If you resend, there are two reservations. So for a timeout on an irreversible tool, you do not retry but hand it to a person to check (if it is an API with an idempotency key, you can safely resend with the same key to confirm — if it is not such an API, you do not). For a rejection (no slot available), you look for open times again and ask anew.
Handing off to a person. When handing off, you keep three rules. If the user asks for a person, you hand off right away at any time. If it fails to understand twice in a row, you do not try a third time. And you summarize the slots already collected and hand them over together — making the receiving person ask everything from the beginning again is the worst kind of handoff.
Test with records, not recordings. When testing an agent, if you compare the wording of speech, the test breaks every time you change the phrasing. You plug in a fake tool with a failure plan and a fake sleep, and compare the state flow and the record of tool calls. For the retry backoff too, you check "how long it tried to pause" without actually waiting.
What it looks like in the field
The LangGraph course treats the same idea as a graph — nodes change state, conditional edges decide what comes next, and it stops where human approval is needed. The MCP server course keeps the same boundary on the tool side — a tool narrows its own permissions and demands confirmation for destructive operations. A voice agent sits between the two, adding the constraint that a person confirms by ear alone, without a screen.
A rule-based parser is not perfect either. This course's parser reads 'I need to see the doctor on Thursday afternoon' as information (inform) rather than a reservation (book) — because the word meaning "reserve" is not in the list. Even so, no major accident occurs inside the state machine. The date is filled, so it asks for the next empty slot (time) and eventually goes through the confirmation question. Even if the understanding module is wrong, irreversible things stop at the confirmation — this is why you use a state machine. Even if you replace the understanding module with an LLM, you leave this structure as it is. Only the component that fills the slots changes.
What you will do in the next lab
You inject failures into a fake clinic reservation API and write down the nature of the tools. You measure the intent accuracy of the LLM and of the rule-based parser. You build the reservation agent as a state machine and add a capability at each step — the happy flow, the "no" to confirmation, retrying a read tool, handing off to a person, and the failure of an irreversible tool. The grader runs nine scenarios directly with your agent.py.