LabHub
Get started
Learn Learning paths Courses

LLM Serving

KV Caches That Don't Fit on the GPU: Cyclic Access, LRU, and Offloading to CPU

Continue in LabHub

In one line

If you read more documents in turn than fit in the GPU's KV cache capacity, LRU, which evicts the least recently used item first, drives the hits to 0 when you read in the same order again. If you offload the evicted KV to CPU memory (the idea behind LMCache), you can load it back instead of recomputing it, and in this measurement the time to first token dropped from 576 ms to 48.6 ms (about 11.9x). The price is that the first read gets about 4% slower.

Why this was needed

vLLM keeps the KV cache of requests that share the same leading part (the prefix) on the GPU and reuses it (prefix caching). If a request puts a long document first and changes only the question, the computation for the document part can be skipped. The problem is that GPU memory is small. In this measurement, after loading a 1.5B model, the remaining KV cache capacity was only 21,520 tokens (GPU KV cache size in the server log), and since one document is 5,011 tokens, a little over four fit. With ten documents the total is 50,110 tokens, about 2.3 times the capacity.

When there is not enough room, something has to be evicted, and this cache evicts the least recently used item first (LRU). The rule assumes that what was used recently will be used again soon, and cyclic access, where you read documents 0 through 9 and then read again from document 0 in the same order, is exactly the case where that assumption is reversed. When the first pass ends, what remains is documents 6–9, the most recently read. In the second pass document 0 is already gone, so it is computed and inserted, which evicts the oldest, document 6. Inserting document 1 evicts document 7. The order of the second read is the same as the order of eviction, so the oldest item is always the one needed soonest. That is why hits drop to 0 even when the capacity falls only slightly short. In the measurement, too, vLLM's prefix_cache_hits was 0 in the second pass.

How it works

LMCache offloads the evicted KV to CPU memory (disk or Redis also work) and, when the same prefix comes again, loads it instead of computing it. The simulator in this lab reduces that flow to three rules.

  1. A GPU hit ends there. The CPU is not touched.
  2. If the GPU misses but the CPU has it, load it from the CPU to the GPU without computing.
  3. If neither has it, compute it and store it in both the GPU and the CPU.

Rule 1 is backed by the measurement log. In the third pass, which reads in reverse, the prefix that vLLM resolved from the GPU was 21,488 tokens, and the number of tokens queried from LMCache was 28,755. That is the total of 50,243 minus the GPU hits. Rule 3 is a simplification of the actual behavior. The report interprets the first read being about 4% slower as the cost of offloading the KV to the CPU while reading, and notes that the KV of ten documents (50,000 tokens) is about 1.3GB, so the configured limit of 1.5GB held it. The CPU side also has a limit, so if you read more documents than the limit allows, the oldest are discarded on the CPU too.

Measured results

Ten documents were read in three passes: a first read, a replay in the same order, and a reverse read. Requests were sent one at a time, and the time to first token (TTFT, ms) was measured. The values are p50 and the mean.

Setup Pass p50 Mean
vLLM only First 575.5 581.2
vLLM only Same order again 576.2 575.9
vLLM only Reverse 505.4 343.6
LMCache First 597.8 605.4
LMCache Same order again 48.6 49.6
LMCache Reverse 47.2 40.9

When reading in the same order again, p50 got about 11.9 times faster, from 576 to 48.6 ms, and the mean got about 11.6 times faster. The time to finish one pass went from 6.2 seconds to 0.9 seconds. On the first read, p50 rose about 4%, from 575.5 to 597.8 ms. In the vLLM-only setup, the second pass is effectively the same as the first (576 versus 575). That means the cache helped not at all.

In the LMCache log, one request in the second pass loaded 4,864 tokens out of about 5,000 (19 multiples of 256) from the CPU and computed only the remaining roughly 150 tokens anew. Loading took 8–9 ms, and computing the same amount anew takes about 560 ms. That loading is more than 60 times cheaper than computing is the whole story of this gap.

The simulator also predicts the reverse-reading pass. With vLLM only, the four documents just read (documents 9, 8, 7 and 6; the numbers are the doc values in the measurement file) hit on the GPU at about 30 ms, and the rest take about 576 ms. With LMCache on, the same four documents take about 30 ms, and the other six are loaded from the CPU in about 46–49 ms. However, in the vLLM-only run, document 5 came out at 438 ms, a value between a hit and a miss. A simulator that puts documents in and takes them out whole cannot draw such a partial hit.

How far to trust these numbers

What it looks like in the field

This measurement did not change LabHub's production llm deployment. To apply it you would have to change the image, the server arguments and the environment variables, and that is a separate task. When you review it, the condition the report points to is that the total tokens of the documents must be greater than GPU KV cache size in the server log for the gap to show. Equally, if the same document is not read again, there is nothing to load. The capacity of the CPU tier has to be decided by first calculating the KV size per token. The 1.5GB in this measurement worked because it held ten documents' worth (about 1.3GB).

What you will do in the next lab

You confirm this flow yourself in pure Python, without a GPU. You build an LRU cache, read 10 documents twice in the same order with a capacity of 21,520 tokens, and see that the hits are 0. You look at which documents hit when you read in reverse, and find how far you have to increase the capacity before hits appear. Then you build a two-tier GPU and CPU cache and confirm that the second pass hits entirely, and predict what happens when the CPU capacity falls short. Finally, you read the per-request measurement file that the table above came from and compute the p50, the mean and the ratio yourself.