LabHub
Get started
Learn Learning paths Courses

Voice AI Agents — a pipeline that listens, looks things up and speaks

A booking agent as a state machine — confirmation, retries, handoff

Continue in LabHub

Goal

Build a phone reservation agent as a state machine so that it never books without a confirmation question, retries only timeouts of read tools, and hands off with a summary when it cannot understand or when the user asks for a person.

Why it matters

Accidents in a voice agent happen in the tools. An irreversible action happens without confirmation, or a timeout is resent and there are two reservations. The tool in this lab is voicekit.clinic.ClinicAPI — a fake reservation API into which you can put failures as a plan — and the user's words are turned into slots by voicekit.nlu.parse. The grader imports VoiceAgent from your /root/voice/agent/agent.py, runs nine scenarios (failure plan + user utterances) directly, and judges not by the wording of speech but by the state and the record of tool calls.

Steps

  1. Inject failures into ClinicAPI and write, for each tool, whether it is irreversible and which errors are worth retrying to /root/voice/agent/tools.json.
  2. Extract the intents of 8 sentences with the LLM (json_schema) and with the rule-based parser, and create /root/voice/agent/intents.jsonl and the correct counts /root/voice/agent/intent_acc.json.
  3. Create /root/voice/agent/agent.py with VoiceAgent(api, sleep) and handle(text), so that it books only after receiving a "yes" to the confirmation.
  4. Handle a "no" to the confirmation question (if there is a new time, confirm again with that time).
  5. On a timeout of the read tool (find_slots), call again up to twice, pausing 0.5 and 1.0 seconds, and if it still fails, hand off with tool_failure.
  6. If the user asks for a person, hand off with user_request, and if it fails to understand twice in a row, hand off with not_understood, and summarize the filled slots when handing off.
  7. For a rejection from the reservation tool (no slot available), look for open times again and ask, and for a timeout, do not resend but hand off.
  8. Leave the turn records of the nine scenarios in /root/voice/agent/traces.jsonl.

Notes

Write down the nature of the tools

Trigger failures yourself with TOOLS and ClinicAPI(plan={"find_slots": ["timeout", "ok"], "book": ["taken"]}) from voicekit.clinic, and write to /root/voice/agent/tools.json tools ({"irreversible": true/false} for each tool name) and retryable (the list of names of errors worth retrying).

The values of TOOLS are (whether it is irreversible, description). ToolTimeout means "there was no response" and ToolError means "it was rejected" — which of the two is worth sending again?

LLM versus rules: who gets the intent right

For 8 sentences — "I'd like to book an appointment for Tuesday" (book), "CAN I CANCEL MY VISIT ON FRIDAY" (cancel), "YEAH THAT WORKS" (yes), "NO THAT'S NOT RIGHT" (no), "CAN I TALK TO A REAL PERSON" (human), "WHAT TIME DO YOU CLOSE TODAY" (hours), "TEN THIRTY IN THE MORNING" (inform), "MY DOG ATE THE REMOTE" (unknown) — in this order, write the result of asking the LLM with a json_schema that binds intent to an enum of the eight values (llm) and the result of voicekit.nlu.parse (rule) to /root/voice/agent/intents.jsonl as {"text", "llm", "rule"}, and write the number that matched the ground truth in parentheses to /root/voice/agent/intent_acc.json as {"llm": n, "rule": m}.

After voice-llm up, voicekit.llm.chat(messages, max_tokens=40, json_schema=schema). The grammar enforces the format, but the content is up to the model. If the result looks odd, that is exactly what this step wants you to see.

Book only after receiving confirmation

In /root/voice/agent/agent.py, create VoiceAgent(api, sleep=time.sleep, max_retries=2, backoff=(0.5, 1.0)) and handle(text). Fill the slots (day, time, name) with parse, ask for the empty slots in the order date → time (finding open times with find_slots and offering them) → name, read back at CONFIRM, and call api.book(day, time, name) only when you hear a "yes" (intent yes), so that it reaches DONE. Grading: in a five-turn conversation and in a two-turn conversation where everything is said at once, book is called exactly once on the last turn.

If you have a single function (_next) that decides "what to ask next," the case where several slots are said at once is handled by the same code. If the time is not among the open times offered (offered), offer them again. If the parser cannot extract a name on the name turn, use the last word of the utterance as the name.

A "no" to confirmation is a correction

When intent is no at CONFIRM: if the utterance has a new time or date, change only that slot and confirm again (CONFIRM), and if it has no information at all, clear the time and offer open times again. Either way, do not call book. Grading: in '…ten a m…' → 'no, make it two thirty' → 'yes', book is called once on the last turn with 14:30.

parse('no, make it two thirty') gives both intent no and time 14:30. If you unconditionally go back to the start on a no, you throw away the information the user just gave.

Retry a timeout of a read tool

If find_slots raises ToolTimeout, call it again up to max_retries (2) times, pausing 0.5 seconds and then 1.0 seconds with self.sleep(backoff[시도 번호]) (the placeholder is the attempt number). If all three fail, set the state to HANDOFF and handoff.reason to tool_failure. Grading: if only the first call times out, it is called twice and, with slept times [0.5], reaches ASK_TIME; if all three fail, it is called three times and, with slept times [0.5, 1.0], reaches HANDOFF.

You must accept sleep as a constructor argument so that the grader can record "how long it tried to pause" without waiting. Making it testable is also design.

Hand off to a person — with a summary

In any state, if intent is human, go to HANDOFF (user_request) immediately. If intent is unknown, ask the user to repeat once, and if it is the second time in a row, go to HANDOFF (not_understood). If it understands, reset the count to 0. When handing off, put the day, time, and name filled so far in handoff.summary. Grading: on 'hmm' → 'blah blah' it hands off on the second turn, and on 'Can I book on Friday' → 'can I talk to a person' the summary has friday.

On the name turn (ASK_NAME), even if the parser returns unknown, the person may have simply said a name — in that case, count it as understood.

When an irreversible tool fails

If book raises ToolError (the slot is filled), do not resend; clear the time and look for open times again with find_slots and offer them (ASK_TIME). If book raises ToolTimeout, do not resend and hand off with HANDOFF (tool_failure). Grading: after a no-slot rejection, get confirmation again with the new time 14:30 and book is called twice (10:00 fails, 14:30 succeeds); on a reservation timeout, book is called once and then HANDOFF.

A timeout on a reservation is not "it did not work" but "we do not know." The server may have created the reservation and only the response was late — if you resend, there are two reservations.

The turn records of the nine scenarios

Run the nine scenarios (happy · one_shot · confirm_no · retry_ok · retry_fail · not_understood · wants_human · taken · book_timeout — the failure plans and user utterances are the same as in the grading descriptions of each step's task) with your agent.py, and write {"scenario", "turn", "user", "state", "say", "tools"(이번 턴에 부른 도구 이름), "outcomes"} (the placeholder in tools is the names of the tools called this turn) for each turn to /root/voice/agent/traces.jsonl. The grader runs the same scenarios again and cross-checks.

Tool records accumulate in api.calls as (name, arguments, result). Extract only what was called this turn from the difference in length before and after the turn. This record is the agent's regression test material.