KV Caches That Don't Fit on the GPU: Cyclic Access, LRU, and Offloading to CPU
In one line
If you read more documents in turn than fit in the GPU's KV cache capacity, LRU, which evicts the least recently used item first, drives the hits to 0 when you read in the same order again. If you offload the evicted KV to CPU memory (the idea behind LMCache), you can load it back instead of recomputing it, and in this measurement the time to first token dropped from 576 ms to 48.6 ms (about 11.9x). The price is that the first read gets about 4% slower.
Why this was needed
vLLM keeps the KV cache of requests that share the same leading part (the prefix) on the GPU and reuses it (prefix caching). If a request puts a long document first and changes only the question, the computation for the document part can be skipped. The problem is that GPU memory is small. In this measurement, after loading a 1.5B model, the remaining KV cache capacity was only 21,520 tokens (GPU KV cache size in the server log), and since one document is 5,011 tokens, a little over four fit. With ten documents the total is 50,110 tokens, about 2.3 times the capacity.
When there is not enough room, something has to be evicted, and this cache evicts the least recently used item first (LRU). The rule assumes that what was used recently will be used again soon, and cyclic access, where you read documents 0 through 9 and then read again from document 0 in the same order, is exactly the case where that assumption is reversed. When the first pass ends, what remains is documents 6–9, the most recently read. In the second pass document 0 is already gone, so it is computed and inserted, which evicts the oldest, document 6. Inserting document 1 evicts document 7. The order of the second read is the same as the order of eviction, so the oldest item is always the one needed soonest. That is why hits drop to 0 even when the capacity falls only slightly short. In the measurement, too, vLLM's prefix_cache_hits was 0 in the second pass.
How it works
LMCache offloads the evicted KV to CPU memory (disk or Redis also work) and, when the same prefix comes again, loads it instead of computing it. The simulator in this lab reduces that flow to three rules.
- A GPU hit ends there. The CPU is not touched.
- If the GPU misses but the CPU has it, load it from the CPU to the GPU without computing.
- If neither has it, compute it and store it in both the GPU and the CPU.
Rule 1 is backed by the measurement log. In the third pass, which reads in reverse, the prefix that vLLM resolved from the GPU was 21,488 tokens, and the number of tokens queried from LMCache was 28,755. That is the total of 50,243 minus the GPU hits. Rule 3 is a simplification of the actual behavior. The report interprets the first read being about 4% slower as the cost of offloading the KV to the CPU while reading, and notes that the KV of ten documents (50,000 tokens) is about 1.3GB, so the configured limit of 1.5GB held it. The CPU side also has a limit, so if you read more documents than the limit allows, the oldest are discarded on the CPU too.
Measured results
Ten documents were read in three passes: a first read, a replay in the same order, and a reverse read. Requests were sent one at a time, and the time to first token (TTFT, ms) was measured. The values are p50 and the mean.
| Setup | Pass | p50 | Mean |
|---|---|---|---|
| vLLM only | First | 575.5 | 581.2 |
| vLLM only | Same order again | 576.2 | 575.9 |
| vLLM only | Reverse | 505.4 | 343.6 |
| LMCache | First | 597.8 | 605.4 |
| LMCache | Same order again | 48.6 | 49.6 |
| LMCache | Reverse | 47.2 | 40.9 |
When reading in the same order again, p50 got about 11.9 times faster, from 576 to 48.6 ms, and the mean got about 11.6 times faster. The time to finish one pass went from 6.2 seconds to 0.9 seconds. On the first read, p50 rose about 4%, from 575.5 to 597.8 ms. In the vLLM-only setup, the second pass is effectively the same as the first (576 versus 575). That means the cache helped not at all.
In the LMCache log, one request in the second pass loaded 4,864 tokens out of about 5,000 (19 multiples of 256) from the CPU and computed only the remaining roughly 150 tokens anew. Loading took 8–9 ms, and computing the same amount anew takes about 560 ms. That loading is more than 60 times cheaper than computing is the whole story of this gap.
The simulator also predicts the reverse-reading pass. With vLLM only, the four documents just read (documents 9, 8, 7 and 6; the numbers are the doc values in the measurement file) hit on the GPU at about 30 ms, and the rest take about 576 ms. With LMCache on, the same four documents take about 30 ms, and the other six are loaded from the CPU in about 46–49 ms. However, in the vLLM-only run, document 5 came out at 438 ms, a value between a hit and a miss. A simulator that puts documents in and takes them out whole cannot draw such a partial hit.
How far to trust these numbers
- This scenario is the most favorable to LMCache. The same documents are read again, the prefix is long at 5,000 tokens, and the GPU cache is short. The report interprets the gain as depending on the share of long prefixes that get reused, but this time that share was not varied and measured.
- Requests were sent one at a time (concurrency 1). This measurement does not show what happens when requests overlap.
- The model is small at 1.5B, and there is a single GPU. The gap is expected to be larger with a large model, but this measurement did not confirm that.
- The documents are synthetic, made of random words, so unlike real documents their prefixes do not partially overlap.
- It is a single measurement, and the standard deviation was not computed. Even if the direction is certain, it is better not to trust digits after the decimal point.
What it looks like in the field
This measurement did not change LabHub's production llm deployment. To apply it you would have to change the image, the server arguments and the environment variables, and that is a separate task. When you review it, the condition the report points to is that the total tokens of the documents must be greater than GPU KV cache size in the server log for the gap to show. Equally, if the same document is not read again, there is nothing to load. The capacity of the CPU tier has to be decided by first calculating the KV size per token. The 1.5GB in this measurement worked because it held ten documents' worth (about 1.3GB).
What you will do in the next lab
You confirm this flow yourself in pure Python, without a GPU. You build an LRU cache, read 10 documents twice in the same order with a capacity of 21,520 tokens, and see that the hits are 0. You look at which documents hit when you read in reverse, and find how far you have to increase the capacity before hits appear. Then you build a two-tier GPU and CPU cache and confirm that the second pass hits entirely, and predict what happens when the CPU capacity falls short. Finally, you read the per-request measurement file that the table above came from and compute the p50, the mean and the ratio yourself.