首个 token 何时到来 — 手动读取 SSE,测量 prefill·缓存·取消
目标
亲自读取 Pod 内小型 LLM 服务器的流式响应,测量第一个 token 的延迟、prefill、KV 缓存和第一句话的时刻,并通过断开连接让生成停止。
为什么重要
语音助手在第一句话结束的那一刻就可以开始说话。让这个时刻变晚的最大部分是提示词计算(prefill),而减少它的最便宜的办法是 KV 缓存。两者都要用数字来看,才能判断提示词该怎么写。评分器不会询问 LLM 服务器,只读取你保存的原始 SSE、测量 JSON 和服务器日志,并且不看因机器而异的绝对时间,而是看同一次运行内部的关系(第一个 token < 总时间,长提示词 > 短提示词,缓存命中 < 无缓存)。
步骤
- 用
voice-llm up启动服务器,把/health保存到/root/voice/llm/health.json,把/props保存到/root/voice/llm/props.json。 - 创建逐行读取 SSE 的
/root/voice/llm/sse.py,把收到的行原样保存到/root/voice/llm/sse_raw.txt,把带时刻的内容片段保存到/root/voice/llm/chunks.jsonl。 - 用
cache_prompt: false把同一个问题发送五次,把第一个 token、总时间和 token 数写入/root/voice/llm/ttft.json。 - 把知识库(KB)文档 0 · 4 · 12 份附加到系统提示词(关闭缓存),把 prompt_n 和第一个 token 的时刻写入
/root/voice/llm/prefill.json。 - 开启缓存,把同一个长提示词发送两次,再在最前面加上一行时间发送一次,写入
/root/voice/llm/cache.json。 - 以流式接收三句话的回答,把第一句话结束的时刻写入
/root/voice/llm/sentence.json。 - 请求一个长回答后,在 0.5 秒时断开连接,把服务器停止所需的时间写入
/root/voice/llm/cancel.json。 - 创建汇总了测量结果的
/root/voice/llm/report.json。
参考
- 服务器:
voice-llm up | status | log | down。地址http://127.0.0.1:8080,模型 Qwen2.5-0.5B-Instruct Q4_0,上下文 4096,slot 1 个,--cache-ram 0。 - 请求正文:
{"messages": [...], "stream": true, "max_tokens": 64, "temperature": 0, "seed": 7, "cache_prompt": false}。用标准库http.client发送,并用resp.readline()读取,就能随片段到来而看到它们。 - 常见错误:把只含 role 的第一个片段当作第一个 token,测量 prefill 时开着缓存,说要断开却只是停止读取(必须关闭连接,服务器才会停止)。
- 文档:llama.cpp server README · Server-Sent Events(HTML 标准) · Qwen2.5-0.5B-Instruct 模型卡
启动服务器,看看启动了什么
用 voice-llm up 启动 LLM 服务器后,把 curl -s localhost:8080/health 保存到 /root/voice/llm/health.json,把 curl -s localhost:8080/props 保存到 /root/voice/llm/props.json。
如果 /health 是 {"status":"ok"},就说明模型已经加载完毕。/props 中有模型路径、上下文长度(default_generation_settings.n_ctx)和 slot 数量(total_slots)。
逐行读取 SSE
在 /root/voice/llm/sse.py 中创建 stream(messages, max_tokens=64, cache_prompt=True, raw_out=None),以 stream: true 发送到 /v1/chat/completions,把收到的行原样写入 raw_out,并为每个有内容的片段收集 {"t_ms": 요청 직전부터의 ms, "text": 조각}(占位符依次为自请求前一刻起的 ms、片段),返回(片段列表、总 ms、timings)。用系统提示词“You are a clinic phone assistant. Answer in one short sentence.”和问题“What should a new patient bring?”运行一次,生成 /root/voice/llm/sse_raw.txt 和 /root/voice/llm/chunks.jsonl。
如果行以“data: ”开头,就把后面作为 JSON 解开。“data: [DONE]”是结束。choices[0].delta.content 为空的片段(第一个片段的 role)要跳过。时刻用 time.perf_counter() 来测量。
测量五次第一个 token 的延迟
用 cache_prompt=False 把 02 的问题发送五次,生成 /root/voice/llm/ttft.json:在 runs 中写入每次的 ttft_ms(第一个内容片段的时刻)、total_ms、tokens(timings 的 predicted_n),在 ttft_p50、total_p50 中写入这两个时间的中位数。
关闭缓存后,每次都会重新计算整个提示词。中位数是排序后五个值中的第三个。第一次可能因为服务器刚启动而稍慢——所以要测量多次,而不是一次。
提示词一长,第一个 token 就会变晚
按名称顺序读取 /opt/lab/fixtures/voice/kb/*.md,从前面取 0 · 4 · 12 份附加到系统提示词,用 cache_prompt=False、max_tokens=32 发送问题“When is the clinic open on Saturday?”,把三条 {"docs": n, "prompt_n": …, "prompt_ms": …, "ttft_ms": …} 作为数组写入 /root/voice/llm/prefill.json。
prompt_n、prompt_ms 在最后一个片段的 timings 中。请观察,随着 token 增多,第一个 token 几乎成比例地变晚。如果缓存开着,与前一次请求重叠的前半部分会被跳过,数字就会变得模糊。
KV 缓存——相同的前半部分会被跳过
使用附加了全部 KB 文档的长系统提示词,先发送一个较短的其他请求(“hi”)把 slot 清空,然后用 cache_prompt=True 发送两次(first、second),再在系统提示词最前面加上“Current time: 14:05. ”再发送一次(changed_prefix),把各自的 prompt_n、ttft_ms 写入 /root/voice/llm/cache.json。
服务器会把 slot 中保留的上一次请求的 K·V 与新请求从前往后比较,跳过相同的部分。如果最前面一行不同,从第一个 token 起就不同,就什么也无法复用。也想一想,如果把会变化的行加到最后面,会怎样。
第一句话什么时候结束
用系统提示词“You are a clinic phone assistant. Answer in exactly three short sentences.”和问题“How do I prepare for a fasting blood test?”以 max_tokens=120 进行流式接收,把拼接起来的文字中第一次在 [.!?] 之后出现空格的那个片段的时刻写作 first_sentence_ms,把那句话写作 first_sentence,生成 /root/voice/llm/sentence.json(ttft_ms、first_sentence_ms、total_ms、first_sentence、text)。如果直到最后都没有出现空格,整段就是一句话。
只看句号就截断,在“3.5”或“a.m.”处会出错。符号后面跟着空格的那一刻,是“句子结束了”的更安全的信号。小模型经常违反“三句话”的指令——所以要用代码来处理长度。
被打断时就停止生成
以 max_tokens=256 流式接收“List the numbers from 1 to 300, separated by commas.”,在 0.5 秒时关闭连接,把到那时为止收到的内容片段数写作 received_tokens,把关闭之后 /slots 的 is_processing 变为 false 为止的 ms 写作 slot_idle_after_ms,写入 /root/voice/llm/cancel.json(max_tokens、received_tokens、closed_at_ms、slot_idle_after_ms)。
http.client 连接的 close() 就是“停”。如果服务器日志(voice-llm log)中打印出“cancel task”,就说明服务器听懂了。请选择一个肯定要说很久的请求——如果在断开之前生成就结束了,就起不到测试的作用。
报告
在 /root/voice/llm/report.json 中,抄写 ttft_p50_ms、total_p50_ms(ttft.json),ttft_12docs_ms、prompt_n_12docs(prefill.json 的最后一项),cache_ttft_ms(cache.json 的 second),first_sentence_ms(sentence.json),cancel_idle_ms(cancel.json 的 slot_idle_after_ms)。
读取前面步骤的文件抄写过来。把数字并排放在一起,就能看出该缩短哪里才能更快开口。