Checking KV Cache Offload with LRU and Two-Tier Cache Simulators
Goal
Use a simulator you build yourself to see how an LRU cache breaks down when you read more documents in turn than fit in the GPU's KV cache capacity, and see how things change when you add a CPU tier. At the end, you compute the ratio yourself from the per-request file that the real measurement left behind. You use only Python, not the GPU.
Why it matters
When a request puts a long document first and changes only the question, reusing the document's KV cache is the biggest saving. But the GPU has little room for that cache, so when there are more documents than capacity, the oldest are evicted first. The first finding of this lab is that when this eviction meets cyclic access, the hits drop to 0 even when the capacity falls only slightly short. The second finding is that if you offload what gets evicted to CPU memory, you can load it back instead of recomputing it, and in the 2026-10-03 measurement that difference was 576 ms versus 48.6 ms to the first token. The third is the limits. The simulator works in whole documents, and the measurement was concurrency 1, a 1.5B model, a single GPU and synthetic documents. The CPU capacity of 56000 tokens in the lab is an approximation: the measurement's CPU limit of 1.5GB divided by the KV size per document (0.1299GB for 4,864 tokens). Step 7 (when CPU capacity falls short) is a simulator prediction, not something that was measured.
Steps
- In
/root/kvlab/kvsim.py, createLRUCache(capacity).access(key, size)returns True on a hit and moves that entry to most recent. On a miss it returns False, evicts the oldest entries first until there is room, and then stores the new one. An entry larger than the capacity is not stored, and no other entry is evicted.key in cachedoes not change the order,cache.usedis the sum of the stored sizes (an attribute), andcache.keys()is the list of keys from the oldest. - Create
/root/kvlab/scenario.py. The arguments are--gpu(default 21520),--docs(default 10),--tokens(default 5011) and--passes(defaultfwd,fwd,rev). The keys are the document numbers 0 through docs-1, and every size is tokens. Afwdpass reads from 0 in order, and arevpass reads in reverse. Print exactly one JSON object to standard output:{"passes": [{"order": "fwd", "gpu_hits": 정수, "cpu_hits": 0, "misses": 정수, "hit_docs": [적중한 문서 번호를 읽은 순서대로]}, ...]}(one object per pass; the placeholders are an integer for each count and the numbers of the hit documents in the order they were read). Save the output of running the default scenario with--passes fwd,fwdto/root/kvlab/cyclic.json. - Save the output of running with
--passes fwd,fwd,revto/root/kvlab/baseline.json. Read which documents are in thehit_docsof the third pass, and why those documents. - Run with
--passes fwd,fwdwhile changing the GPU capacity among 21520, 30000, 40000, 45000, 50109 and 50110, and save the number of hits in the second pass at each capacity (the sum of gpu_hits and cpu_hits) to/root/kvlab/cliff.jsonas{"자리": 적중 수}(the placeholders are the capacity and the number of hits). The keys are strings. - In
/root/kvlab/kvsim.py, addTwoTierCache(gpu_capacity, cpu_capacity)(leave LRUCache as it is). The attributesgpuandcpuare each an LRUCache, andaccess(key, size)returns one of"gpu","cpu"and"miss". (1) On a GPU hit it returns"gpu"and the CPU is not touched. (2) If the GPU misses but the CPU has it, it returns"cpu", makes the CPU copy the most recent, and loads it into the GPU as well. (3) If both miss, it returns"miss"and the entry is stored in both the GPU and the CPU. - In
scenario.py, add--cpu(default 0) and run it with TwoTierCache. In the output,gpu_hitsis the GPU hits,cpu_hitsis the number loaded from the CPU, andhit_docsis the document numbers of both combined. If--cpuis not given or is 0, the result must be the same as in step 2. Save the output of--gpu 21520 --cpu 56000 --passes fwd,fwd,revto/root/kvlab/offload.json. - Keep the GPU capacity at 21520 and run with
--passes fwd,fwdwhile changing the CPU capacity among 0, 20000, 30000, 40000, 50109, 50110 and 56000. Save the number of misses in the second pass (misses) at each CPU capacity to/root/kvlab/cpu_cliff.jsonas{"CPU 자리": 미적중 수}(the placeholders are the CPU capacity and the number of misses). The keys are strings. - Copy the measured material with
cp /opt/fixtures/kvoffload/ttft.jsonl /root/kvlab/ttft.jsonl(do not modify the contents). Then create/root/kvlab/summarize.py <입력.jsonl> <출력.json>(the placeholders are the input file and the output file). The input is one JSON object per line (config,pass,doc,doc_tokens,ttft_ms,total_ms), and blank lines are skipped. The output is{"baseline": {패스: {"n", "p50_ms", "mean_ms"}}, "lmcache": {패스: {...}}, "ratio": {패스: {"p50", "mean"}}}(the placeholder is the pass name). p50 is the middle value (the mean of the two middle values when the count is even), ms has one decimal place, and ratio is baseline divided by lmcache with two decimal places. The pass names are exactly as they appear in the input (pass1-cold,pass2-replay,pass3-reverse). Run it on the copied file and save the result to/root/kvlab/measured.json.
Notes
- The work files (
/root/kvlab) disappear when the session ends, so copy them beforehand if you need them. - The simulator puts documents in and takes them out whole. In the vLLM-only reverse read in the measurement, document 5 looked like a partial hit at 438 ms, which this simulator cannot draw.
- Common mistake 1: not moving an entry to most recent on a hit. The cache becomes one that evicts in insertion order instead of LRU.
- Common mistake 2: emptying the cache first when you meet an entry larger than the capacity. You lose healthy entries without even being able to store the new one.
- Common mistake 3: creating a new cache for every pass. The cache has to carry over between passes for the second pass to mean anything.
- The grader runs your
kvsim.py,scenario.pyandsummarize.pydirectly on inputs it has not shown you. Writing numbers into the JSON will not pass.
Build an LRU cache
In /root/kvlab/kvsim.py, create LRUCache(capacity). access(key, size) returns True on a hit and moves that entry to most recent. On a miss it returns False, evicts the oldest entries first until there is room, and then stores the new one. An entry larger than the capacity is not stored, and no other entry is evicted. key in cache does not change the order, cache.used is the sum of the stored sizes (an attribute), and cache.keys() is the list of keys from the oldest.
Python's OrderedDict fits this job, since it has move_to_end, which moves an entry to the very end, and popitem(last=False), which takes the first one. If you forget to change the order on a hit, you get FIFO, which evicts in insertion order, instead of LRU. If you empty the cache first when you meet an entry larger than the capacity, you lose healthy entries too.
Read twice in the same order
Create /root/kvlab/scenario.py. The arguments are --gpu (default 21520), --docs (default 10), --tokens (default 5011) and --passes (default fwd,fwd,rev). The keys are the document numbers 0 through docs-1, and every size is tokens. A fwd pass reads from 0 in order, and a rev pass reads in reverse. Print exactly one JSON object to standard output: {"passes": [{"order": "fwd", "gpu_hits": 정수, "cpu_hits": 0, "misses": 정수, "hit_docs": [적중한 문서 번호를 읽은 순서대로]}, ...]} (one object per pass; the placeholders are an integer for each count and the numbers of the hit documents in the order they were read). Save the output of running the default scenario with --passes fwd,fwd to /root/kvlab/cyclic.json.
Use the same single cache across passes (do not empty the cache when the pass changes). Put the hit documents in hit_docs in the order they were read. If you mix explanatory sentences into the output, the grader cannot read it as JSON. You can save it like python3 scenario.py --passes fwd,fwd > cyclic.json.
What remains when you read in reverse
Save the output of running the same scenario with --passes fwd,fwd,rev, so that it reads first, then again in the same order, then in reverse, to /root/kvlab/baseline.json. Read which documents are in the hit_docs of the third pass, and why those documents.
No new code is needed. Run the scenario.py from step 2 with only the arguments changed. Think about what is left in the cache when the third pass starts, and in what order you meet it when you read in reverse.
How far to increase the capacity before hits appear
Run with --passes fwd,fwd while changing the GPU capacity among 21520, 30000, 40000, 45000, 50109 and 50110, and save the number of hits in the second pass at each capacity (the sum of gpu_hits and cpu_hits) to /root/kvlab/cliff.json as {"자리": 적중 수} (the placeholders are the capacity and the number of hits). The keys are strings.
There is a range where hits do not appear even if you nearly double the capacity. Compute the sum of the sizes of the ten documents, then run the capacity that is 1 token short of it and the capacity that matches it exactly, and compare them. Do not run the six cases by hand; use a loop and build the JSON.
A two-tier GPU and CPU cache
In /root/kvlab/kvsim.py, add TwoTierCache(gpu_capacity, cpu_capacity) (leave LRUCache as it is). The attributes gpu and cpu are each an LRUCache, and access(key, size) returns one of "gpu", "cpu" and "miss". There are three rules. (1) On a GPU hit it returns "gpu" and the CPU is not touched. (2) If the GPU misses but the CPU has it, it returns "cpu", makes the CPU copy the most recent, and loads it into the GPU as well. (3) If both miss, it returns "miss" and the entry is stored in both the GPU and the CPU.
One LRUCache per tier is enough. To know whether it hit, you must first ask with in before calling access (on a miss, access goes as far as storing). It must also work when the CPU capacity is 0, and then it becomes the same as a one-tier cache.
What happens to the second pass when you offload to the CPU
In scenario.py, add --cpu (default 0) and run it with TwoTierCache. In the output, gpu_hits is the GPU hits, cpu_hits is the number loaded from the CPU, and hit_docs is the document numbers of both combined. If --cpu is not given or is 0, the result must be the same as in step 2. Save the output of --gpu 21520 --cpu 56000 --passes fwd,fwd,rev to /root/kvlab/offload.json.
Change the cache that run creates from LRUCache to TwoTierCache, and count each of the three return values of access. Run it again without --cpu to check that the scenario from the previous step is not broken. Compare the result with the measurement: in the reverse-reading pass, how many GPU hits and CPU hits are there, and does that match the number of documents that split into about 30 ms and about 48 ms in the measurement?
When the CPU capacity falls short
Keep the GPU capacity at 21520 and run with --passes fwd,fwd while changing the CPU capacity among 0, 20000, 30000, 40000, 50109, 50110 and 56000. Save the number of misses in the second pass (misses) at each CPU capacity to /root/kvlab/cpu_cliff.json as {"CPU 자리": 미적중 수} (the placeholders are the CPU capacity and the number of misses). The keys are strings.
Only the argument you change and the value you read differ from the loop in step 4. Look at the result and judge whether the capacity the CPU tier needs is the share that did not fit in the GPU or the whole set of documents. This value is a simulator prediction. This measurement only measured the case where the CPU limit held all the documents.
Compute the ratios from the measured file
Copy the measured material with cp /opt/fixtures/kvoffload/ttft.jsonl /root/kvlab/ttft.jsonl (do not modify the contents). Then create /root/kvlab/summarize.py <입력.jsonl> <출력.json> (the placeholders are the input file and the output file). The input is one JSON object per line (config, pass, doc, doc_tokens, ttft_ms, total_ms), and blank lines are skipped. The output is {"baseline": {패스: {"n", "p50_ms", "mean_ms"}}, "lmcache": {패스: {...}}, "ratio": {패스: {"p50", "mean"}}} (the placeholder is the pass name). p50 is the middle value (the mean of the two middle values when the count is even), ms has one decimal place, and ratio is baseline divided by lmcache with two decimal places. The pass names are exactly as they appear in the input (pass1-cold, pass2-replay, pass3-reverse). Run it on the copied file and save the result to /root/kvlab/measured.json.
For each pair of config and pass, collect ttft_ms and use statistics.median and statistics.mean. When reading the result, look at whether ratio is greater or less than 1. What does a value below 1 in the first pass mean? The grader also runs this script with inputs whose number and order of documents differ.