KV Cache Memory: Tokens Left on an H200

KV Cache Memory: How Many Llama 3 Tokens Fit on One H200

A capacity sheet that sizes an H200 from the model name alone will accept a 70-billion-parameter deployment that, at 16-bit weights, has almost no room left for context. KV cache memory for the Llama 3 70B architecture in Table 3 of the Llama 3 paper is 327,680 bytes per token at 2 bytes per element. The same paper’s rounded parameter count, stored at 2 bytes each, is 140 billion bytes. NVIDIA states that the H200 has 141 gigabytes of HBM3e. Subtracting on a decimal-gigabyte reading leaves 1.0 billion bytes, enough for 3,052 tokens of cache in total, not per user.

Thesis: on one H200, the binding constraint for Llama 3 8B is the cache, and the binding constraint for Llama 3 70B and 405B at BF16 or FP16 is the weights. A buyer who treats “141 GB” as plenty for a 70B model is using the wrong residual. The number to put in the purchase order is the residual after weights, divided by bytes per token, with the gigabyte convention written down.

Every byte figure below is either taken from a primary source or computed from the formula next to it. We did not run the models. Parameter counts are the rounded counts the paper publishes (8B, 70B, 405B), not a dump of the checkpoint tensors. The KV cache memory total for a job is that per-token cost times the tokens you actually keep.

The KV cache memory formula, checked against the PagedAttention paper

Kwon, Li, Zhuang, Sheng, Zheng, Yu, Gonzalez, Zhang, and Stoica give a worked token in the PagedAttention paper. For the 13B OPT model, one token’s key and value tensors take 2 × 5,120 × 40 × 2 bytes, which they report as 800 KB, and 2,048 tokens as 1.6 GB. The product 2 × 5,120 × 40 × 2 is exactly 819,200 bytes, which is 800 × 1,024. Their “800 KB” is 800 kibibytes. Across 2,048 tokens the total is 1,677,721,600 bytes: 1.5625 gibibytes, which rounds to the 1.6 GB they print.

The same identity is the sizing formula. Bytes per token equal 2 × layers × key-value hidden size × bytes per element. The leading 2 is one factor for keys and one for values. For multi-head attention the key-value hidden size is the model dimension. For grouped-query attention it is key-value heads times head dimension, which is smaller. Using the full model dimension on a grouped-query model overstates the cache by the ratio of query heads to key-value heads.

The paper also records what a bad allocator does with that formula. On the systems they profiled, only 20.4 percent to 38.2 percent of reserved cache memory held actual token state, because those systems pre-allocated a contiguous chunk out to the maximum length. vLLM’s claim, in the same paper, is near-zero waste by allocating fixed-size blocks on demand. A 2026 capacity check should therefore use the formula itself when the engine is a paged allocator, and should inflate the reservation only when the engine still pre-allocates the maximum length. The formula is not a throughput benchmark. It does not include activations, CUDA graphs, or a framework overhead reserve.

Llama 3’s published shape, and what one token costs

Table 3 of the Llama 3 paper (Dubey and others, arXiv:2407.21783) lists three dense models. All three use grouped-query attention with 8 key-value heads. The 8B model has 32 layers, model dimension 4,096, and 32 attention heads. The 70B model has 80 layers, dimension 8,192, and 64 heads. The 405B model has 126 layers, dimension 16,384, and 128 heads. The paper does not print head dimension as its own row. Head dimension is model dimension divided by attention heads, which is 128 in every column. Key-value hidden size is then 8 × 128 = 1,024 for all three.

The paper states that the flagship model supports a context window of up to 128K tokens, and that Llama 3.1 releases pre-trained and post-trained versions of the 8B, 70B, and 405B models. Table 3 is captioned as Llama 3 hyperparameters and includes the 405B column. We apply it to those three published shapes. We do not apply it to a fine-tune whose config file changes the head count.

At 2 bytes per element, which covers FP16 and BF16, bytes per token are 2 × layers × 1,024 × 2. That is 131,072 bytes for 8B (128 kibibytes), 327,680 bytes for 70B (320 kibibytes), and 516,096 bytes for 405B (504 kibibytes). At 1 byte per element the cache halves. The paper does not state the serving dtype. Two bytes is the comparison case used in the PagedAttention example; one byte is the FP8 case a Blackwell-era stack may actually run. Activations and the KV cache of the current step are extra.

Bar chart comparing KV cache per token for Llama 3 8B, 70B, and 405B. BF16 values are 128, 320, and 504 kibibytes. FP8 values are half of those.
Figure 1. Bytes per token from Table 3 of the Llama 3 paper, with head dimension derived as model dimension divided by attention heads, and 8 key-value heads. Source: arXiv:2407.21783, retrieved October 2, 2026; arithmetic by GPU Insights.
KV cache memory per token for the three Llama 3 shapes in Table 3
Model (published count)LayersKV heads × head dimBytes/token at 2 bytesBytes/token at 1 byte128K context at 2 bytes
8B328 × 128131,07265,53616.78 GB
70B808 × 128327,680163,84041.94 GB
405B1268 × 128516,096258,04866.06 GB

Source: layer and head counts from Table 3, arXiv:2407.21783; head dimension and byte totals computed by GPU Insights, retrieved October 2, 2026. Gigabyte figures are decimal (109 bytes). Takeaway: grouped-query attention makes per-token cache depend on layers, not on the 4,096-to-16,384 model dimension.

What remains on a 141 GB H200 after the weights

NVIDIA’s H200 product page states that the H200 SXM and the H200 NVL each offer 141 GB of memory and 4.8 TB/s of bandwidth, and it labels the specification table preliminary and subject to change. The page uses the word gigabytes. Read as decimal units, 141 GB is 141 × 109 bytes. Memory marketing sometimes uses GB for gibibytes. Both readings are computed below; the decimal reading is the one that matches the word on the page.

Weights at 2 bytes per parameter, using the rounded counts, are 16.0 GB, 140.0 GB, and 810.0 GB. On one H200 the residuals are +125.0 GB, +1.0 GB, and −669.0 GB. The 8B model leaves most of the card for cache. The 70B model consumes 140 / 141 = 99.3 percent of the stated capacity before a single cached token. The 405B model does not fit. At 1 byte per parameter the weights are 8.0, 70.0, and 405.0 GB, and the residuals are +133.0, +71.0, and −264.0 GB. Quantizing the 70B weights is what creates a context budget. Quantizing the 405B weights is not enough for one card.

A 128K context, the maximum the paper states, costs 16.78 GB, 41.94 GB, and 66.06 GB at 2 bytes per element. The 8B model can hold several such contexts in the residual. The 70B model at 2-byte weights cannot hold one. This is a memory-fit statement. It is not a claim about tokens per second, which depends on the engine, the batch, and the latency target. Those comparisons belong in the H200, B200, and H100 cost-per-token analysis and in the MLPerf Inference v6.1 per-GPU recomputation.

Residual memory on one H200 after weights, decimal gigabytes, rounded parameter counts
ModelWeights at 2 bytesResidual at 141 GBTokens that fit in the residualWeights at 1 byteResidual at 1 byte
8B16.0 GB125.0 GB953,6748.0 GB133.0 GB
70B140.0 GB1.0 GB3,05270.0 GB71.0 GB
405B810.0 GBdoes not fit0 on one GPU405.0 GBdoes not fit

Source: NVIDIA H200 memory capacity, product page, retrieved October 2, 2026; parameter counts from arXiv:2407.21783; residuals and token counts computed by GPU Insights. Takeaway: at 2-byte weights the 70B residual is 1.0 decimal GB, or 3,052 tokens.

The gigabyte convention changes the 70B answer

If 141 GB is read as 141 gibibytes, the card holds 141 × 1,0243 = 151,397,597,184 bytes, which is 151.40 decimal gigabytes. The 70B weights at 140 × 109 bytes then leave 11.40 decimal gigabytes, and the token count rises from 3,052 to 34,783. The qualitative conclusion survives: a full 8,192-token sequence is 2.684 GB at 327,680 bytes per token, so the gibibyte reading holds 11.40 / 2.684 = 4.2 such sequences and the decimal reading holds none. The quantitative conclusion does not survive a silent unit choice. Write the unit into the sheet.

The 405B case is insensitive to that ambiguity. Even at 151.40 decimal gigabytes per card, 810 GB of weights need six cards (5 × 151.40 is still under 810). On the decimal reading, five cards hold 705 GB and six hold 846 GB, leaving 36 GB of slack. That slack is 69,754 tokens of 2-byte cache across the whole group, shared by every sequence. At 1-byte weights, 405 GB need three cards (423 GB), leaving 18 GB. Because the byte width and the card count both halve, the slack in tokens is the same 69,754. That equality is an artifact of the rounded 405 billion count. An exact checkpoint a few billion parameters higher or lower moves it.

Home-lab advice does not transfer. The local-hardware guide sizes consumer and workstation cards for a single interactive session. A datacenter H200 serving many sequences spends the residual on the sum of their contexts. One long request and twenty short ones are different problems with the same formula.

Editorial estimate — Methodology: the $0 figure in this section is not a price. No market price is used. Token counts use decimal gigabytes for the primary case. The gibibyte case is a unit sensitivity, not a second NVIDIA specification. No memory is reserved for activations, the framework, or fragmentation. Bytes per token = 2 × layers × 8 × 128 × element bytes.

Worked example. A team wants Llama 3 70B at 2-byte weights on one H200, decimal reading, with an 8,192-token budget per sequence and no prefix sharing. Cache per sequence is 8,192 × 327,680 = 2.684 GB. The residual is 1.0 GB, so the sequence does not fit. The same model at 1 byte per weight leaves 71.0 GB. Dividing by 2.684 GB gives 26.4 sequences of 8,192 tokens before the card is full of weights plus cache, and that count ignores activations. If the engine pre-allocates the 128K maximum instead of 8,192, each sequence reserves 41.94 GB and the same 71.0 GB holds one sequence.

The counterargument: paging, sharing, and quantization already solved this

The serious objection is that production stacks do not store a dense 2-byte cache for every token of every request. PagedAttention allocates blocks on demand. Prefix caching shares a system prompt across requests. FP8 caches, and later 4-bit caches, cut the element width. A sheet that ignores all three will reject deployments that already run.

The objection is right about the direction and wrong about the residual. Paging removes the waste the paper measured at 20.4 to 38.2 percent occupied. It does not shrink the bytes of a token that is actually stored. Sharing helps only when the bytes already exist and a second request can point at them. A cold 8,192-token request with a unique prompt pays the full product. Quantization is the lever that changes the 70B result, and it is a quality decision, not a free bit. The vLLM, SGLang, and TensorRT-LLM comparison is where engine choice belongs. This page stops at the bytes the engine will have to place somewhere.

What this analysis can’t tell you

It cannot tell you the exact parameter count of a checkpoint. “70B” and “405B” are the paper’s rounded figures. A few percent either way moves the 70B residual from a thin positive to a negative on the decimal reading, because 140 is only 1 GB under 141. The 128K column uses 128,000 tokens. If the checkpoint’s maximum position is 131,072, which is 217 and the usual meaning of “128K” in a config file, multiply that column by 131,072 / 128,000 = 1.024. The paper prints “128K,” not 131,072. It cannot tell you serving dtype, activation footprint, or the framework reserve on a given vLLM or TensorRT-LLM version. It cannot tell you H200 capacity if NVIDIA revises the preliminary table. It cannot tell you throughput, time to first token, or cost per million tokens. It does not cover mixture-of-experts models, where the weight-resident set is not the full parameter count, or speculative decoding, which stores extra draft-model state.

The PagedAttention arithmetic was checked against that paper’s OPT-13B example and matches. The Llama arithmetic uses a derived head dimension. If a config file sets a different head_dim or num_key_value_heads, recompute. Do not reuse the 1,024-wide key-value hidden size for a model that is not in Table 3.

When to buy a second GPU, by role

Use one threshold. If 2-byte weights exceed 90 percent of the card’s stated capacity, do not promise multi-thousand-token contexts on that card. On the decimal reading the 70B model is at 99.3 percent, so it fails the threshold. The 8B model is at 11.3 percent, so the cache, not the weights, is what you scale. The 405B model fails on one card at both 1 and 2 bytes per weight.

An ML platform lead who is about to publish an 8K context limit for a 70B model at BF16 should either move the weights to 1 byte per parameter or place the model on two H200s before the limit is announced. A FinOps owner should refuse a single-GPU quote for that shape, because the residual cannot hold the context the product page will advertise. An engineering manager can hand the formula to the person who owns the config file: 2 × layers × key-value heads × head dimension × element bytes × tokens, then subtract weights from 141 × 109 and from 141 × 1,0243. A procurement lead should treat the H200 table as preliminary, the word NVIDIA prints under it, and should not sign a memory commitment that cites 141 GB without the unit.

Where the request should run at all, on a device or in a shared cluster, is a separate decision. The edge-versus-cloud inference framework covers latency and data-residency constraints that this byte count does not.

FAQ

How much KV cache memory does one token use?

For the Llama 3 shapes in Table 3, at 2 bytes per element, one token uses 131,072 bytes (8B), 327,680 bytes (70B), or 516,096 bytes (405B). The formula is 2 × layers × 8 × 128 × element bytes. A model with a different key-value head count needs its own product.

Does a 70B model fit on one H200?

The rounded 70 billion weights at 2 bytes are 140 GB. NVIDIA states 141 GB. On a decimal reading the weights fit and about 3,052 tokens of cache fit with them. On a gibibyte reading about 34,783 tokens fit. An 8,192-token sequence at 2.68 GB fits only under the gibibyte reading, and only a few times. At 1 byte per weight, 71 GB remain for cache.

Why is the 8B cache not 32 times smaller than a model with 32 query heads?

Grouped-query attention stores keys and values for 8 heads, not for every query head. All three Llama 3 shapes in the table use 8 key-value heads and a head dimension of 128, so the per-layer cache is the same width. The 405B model is more expensive per token because it has 126 layers, not because its model dimension is 16,384.

Should I reserve the maximum context for every request?

Only if the allocator does. The PagedAttention paper found that pre-allocating the maximum left 20.4 to 38.2 percent of the reservation occupied by real tokens on the systems they measured. A paged engine charges the tokens that exist. Size the cold, unshared sequence at the length you will actually accept, and add a separate reserve for activations.

Does this predict tokens per second?

No. Memory fit is a condition for the job to start. Tokens per second also depend on bandwidth, batching, the kernel, and the latency caps. Use a measured result, with the hardware and software version attached, for that question.

Sources & further reading

Related reading

Updated: October 2026. This page is an analysis of published specifications and a recomputed byte model. It is not a benchmark, a quote, or advice to buy a particular GPU. H200 specifications are preliminary on NVIDIA’s page. Parameter counts are the rounded figures in the Llama 3 paper.

Iovanny Olguín Ávila
Author: Iovanny Olguín Ávila

Computer Systems Engineer with an MSc in Computer Science. I apply quantitative analysis and data-driven methodologies to evaluate financial instruments, investment vehicles, and emerging technologies. My technical background allows me to cut through marketing language and analyze the actual mechanics of financial products — from HELOC structures to Medicare Advantage plan design to business credit card reward algorithms.

Leave a Comment