Back to all posts
Guide
7 min read

Local LLM VRAM Requirements: How Much Memory You Actually Need

DevToolLab Team

DevToolLab Team

August 15, 2026

Local LLM VRAM Requirements: How Much Memory You Actually Need

Picking hardware for a local LLM usually starts with a rule of thumb, something like half a gigabyte of VRAM per billion parameters, and that rule holds right up until the model loads fine and then runs out of memory partway through a long conversation. The gap between "the model fits" and "the model runs at the context length I need" is where most local setups go wrong.

Two numbers explain it. A Q4_K_M quant of an 8B model is not 4 bits per weight, it is 4.90, so the file is 23 percent larger than the arithmetic suggests. And at 128k context that same model's KV cache is 17.18 GB against 4.92 GB of weights, which is three and a half times the model itself.

Both figures below were computed from a script you can run, checked against the actual file sizes on Hugging Face.

Weights: the Quant Name Is Not the Bit Count

The naive calculation is parameters times bits divided by eight. For Llama 3.1 8B at Q4 that gives 4.01 GB. The file you actually download is 4.92 GB.

The Hugging Face file listing for bartowski/Meta-Llama-3.1-8B-Instruct-GGUF, showing Q4_K_M at 4.92 GB, Q4_K_S at 4.69 GB, Q4_K_L at 5.31 GB and Q5_K_M at 5.73 GB
The Hugging Face file listing for bartowski/Meta-Llama-3.1-8B-Instruct-GGUF, showing Q4_K_M at 4.92 GB, Q4_K_S at 4.69 GB, Q4_K_L at 5.31 GB and Q5_K_M at 5.73 GB

The difference is not rounding. K-quants store per-block scale and minimum values alongside the quantized weights, and the mixed variants keep sensitive tensors, typically the embedding and output layers, at higher precision than the name suggests. The suffix tells you the dominant format, not the average.

JavaScript
// What a local LLM actually costs in memory, checked against real GGUF files.
const GB = 1e9
const pad = (s, n) => String(s).padStart(n)

// Llama 3.1 8B Instruct, from its config.json
const M = { params: 8.03e9, layers: 32, kvHeads: 8, headDim: 128 }

// Real bartowski/Meta-Llama-3.1-8B-Instruct-GGUF file sizes, in bytes.
const REAL = { Q4_K_M: 4.92e9, Q5_K_M: 5.73e9, Q6_K: 6.60e9, Q8_0: 8.54e9 }
const NOMINAL_BITS = { Q4_K_M: 4, Q5_K_M: 5, Q6_K: 6, Q8_0: 8 }

console.log("weights: naive guess vs the file you actually download\n")
console.log("quant      naive    real    error   real bits/weight")
for (const [q, real] of Object.entries(REAL)) {
  const naive = (M.params * NOMINAL_BITS[q]) / 8
  const bits = (real * 8) / M.params
  const err = ((real - naive) / naive) * 100
  console.log(
    `${q.padEnd(9)} ${pad((naive / GB).toFixed(2), 6)}  ${pad((real / GB).toFixed(2), 6)}  ` +
    `${pad("+" + err.toFixed(0) + "%", 6)}   ${bits.toFixed(2)}`,
  )
}

// KV cache: 2 tensors (K and V) per layer, per token.
const kvPerToken = (bytesPerElem) => 2 * M.layers * M.kvHeads * M.headDim * bytesPerElem

console.log(`\nKV cache = ${(kvPerToken(2) / 1024).toFixed(0)} KB per token at fp16\n`)
console.log("context   KV fp16    KV q8    total (Q4_K_M + fp16 KV)")
for (const ctx of [4096, 8192, 32768, 131072]) {
  const kv16 = kvPerToken(2) * ctx
  console.log(
    `${pad(ctx, 7)}  ${pad((kv16 / GB).toFixed(2), 6)} GB  ${pad((kvPerToken(1) * ctx / GB).toFixed(2), 5)} GB  ` +
    `${pad(((REAL.Q4_K_M + kv16) / GB).toFixed(2), 8)} GB`,
  )
}
text
weights: naive guess vs the file you actually download

quant      naive    real    error   real bits/weight
Q4_K_M      4.01    4.92    +23%   4.90
Q5_K_M      5.02    5.73    +14%   5.71
Q6_K        6.02    6.60    +10%   6.58
Q8_0        8.03    8.54     +6%   8.51

The overhead shrinks as precision rises, from 23 percent at Q4_K_M down to 6 percent at Q8_0, because the block metadata is a fixed cost spread over more bits of payload. If you want one number to plan with, use 0.61 GB per billion parameters at Q4_K_M, which is what 4.90 bits works out to. That is where the familiar "0.6 GB per billion" rule comes from, and now you know why it is 0.6 rather than 0.5.

KV Cache: the Part That Actually Kills You

Weights are a fixed cost you pay once. The KV cache grows linearly with context length, and every token you feed the model adds to it permanently for that conversation.

The formula is two tensors, K and V, for every layer, for every token:

2 x layers x kv_heads x head_dim x bytes_per_element

For Llama 3.1 8B that is 2 x 32 x 8 x 128 x 2 bytes at fp16, which comes to 128 KB per token. Note kv_heads is 8, not 32. Grouped-query attention shares key and value projections across query heads, and without it this model's cache would be four times larger.

text
KV cache = 128 KB per token at fp16

context   KV fp16    KV q8    total (Q4_K_M + fp16 KV)
   4096    0.54 GB   0.27 GB      5.46 GB
   8192    1.07 GB   0.54 GB      5.99 GB
  32768    4.29 GB   2.15 GB      9.21 GB
 131072   17.18 GB   8.59 GB     22.10 GB

Read the last row again. The model weighs 4.92 GB and the cache for its advertised 128k context weighs 17.18 GB. An 8B model, the size everyone calls "runs on anything", needs 22.10 GB to use the context window it ships with. That does not fit on a 16 GB card and leaves nothing spare on a 24 GB one.

This is why a setup that loads fine dies mid-conversation. You sized for the weights.

Serving More Than One Person Multiplies It

Everything above assumes a single conversation. The moment you put a local model behind an API that two people hit at once, the KV cache multiplies by the number of sequences held in flight, because each one keeps its own keys and values.

Four concurrent 8k sessions on Llama 3.1 8B is 4 x 1.07 GB, so 4.29 GB of cache against 4.92 GB of weights. Sixteen concurrent sessions is 17.2 GB and you are out of memory on a 24 GB card while every individual conversation looks modest. This is the difference between running a model for yourself and serving one, and it is why inference servers such as vLLM spend so much engineering on paged attention and cache reuse: the naive allocation wastes memory reserving space for tokens nobody generated yet.

If you are sizing for a team rather than a laptop, multiply the cache column by peak concurrency, not average, and quantize the cache before you do anything else.

What This Means for Real Hardware

VRAMComfortable at Q4_K_MRealistic context
8 GB7B to 8B8k to 16k
12 GB8B to 13B16k to 32k
16 GB13B to 14B32k, or 8B at 64k
24 GB24B to 32B32k comfortably, 8B at full 128k
48 GB70B at Q432k
96 GB+70B at Q6 or abovelong context on large models

Apple Silicon changes the arithmetic because memory is unified rather than dedicated, so a 32 GB M-series machine can hold a model and cache that would need a 32 GB discrete GPU. It is slower per token than an equivalent NVIDIA card, but the ceiling is set by total RAM rather than by what fits on a board.

Four Ways to Buy Headroom

Quantize the KV cache. The q8 column above is the single biggest lever available: half the cache memory for a quality cost most people cannot detect in chat. llama.cpp exposes this as --cache-type-k and --cache-type-v, and Ollama has an equivalent setting. Halving 17.18 GB to 8.59 GB is the difference between fitting on a 24 GB card and not.

Set context deliberately. Runtimes often allocate the full advertised window whether you use it or not. Dropping from 131072 to 32768 on this model frees 12.9 GB. Most chat and coding sessions never approach 32k.

Drop a quant level before dropping a model size. A 13B at Q4_K_M usually beats an 8B at Q8_0 at similar memory. Bigger model, cheaper weights, is generally the better trade until you reach Q3, where quality starts to visibly degrade.

Offload layers rather than giving up. llama.cpp splits layers between GPU and CPU. Performance falls off, sometimes sharply, but a model that runs slowly is more useful than one that does not load.

Measure Instead of Guessing

The estimate tells you what to buy. Once the model is loaded, check the real number rather than trusting the arithmetic, because runtimes add allocator overhead, CUDA context and fragmentation on top of what you calculated.

On an NVIDIA card, nvidia-smi --query-gpu=memory.used,memory.total --format=csv prints used against total, and running it while a long conversation grows shows the cache filling in real time. On Apple Silicon the equivalent view is the memory tab in Activity Monitor, watching for the point where the machine starts swapping, which shows up as tokens per second collapsing rather than as an error.

The signature of a cache problem is specific and worth recognizing: generation is fine for the first few exchanges and then slows sharply or crashes as the conversation lengthens. Weights are allocated once at load, so a failure that arrives later is almost always the cache growing into memory you did not budget.

Sizing a Model You Have Not Downloaded

The script above is specific to Llama 3.1 8B, and adapting it is a matter of reading four numbers out of the model's config.json on Hugging Face: num_hidden_layers, num_key_value_heads, hidden_size and num_attention_heads. Head dimension is hidden_size divided by num_attention_heads.

Watch num_key_value_heads in particular. When it equals num_attention_heads the model uses full multi-head attention and its cache is several times larger per token than a grouped-query model of the same parameter count. Two 8B models can differ by 4x on cache memory for that reason alone, which no parameter-count rule of thumb will ever tell you.

For weights, multiply parameters by the real bits per weight from the table above rather than the number in the quant name, or just check the file size on the repo, which is the ground truth this whole post is calibrated against.

Conclusion

Sizing a local LLM is two calculations, not one. Weights cost parameters times real bits per weight, which for Q4_K_M is about 4.9 rather than 4, giving roughly 0.61 GB per billion parameters. Cache costs 2 x layers x kv_heads x head_dim x bytes per token, and at long context it can dwarf the weights entirely.

The practical takeaway is that context length, not parameter count, is what decides whether a model fits. An 8B at 8k needs 6 GB and the same model at 128k needs 22 GB. Before buying a card or blaming a runtime, run the numbers for the context you actually intend to use, and quantize the cache before you compromise on the model.

Model files and quantization formats change. Check the file size on the repository before committing to hardware.

Related Posts

Playwright vs Cypress vs Selenium in 2026

Playwright, Cypress and Selenium compared on stars and downloads, plus what BrowserStack, Sauce Labs and TestMu AI charge to run them in CI.

By DevToolLab Team•

Best API Documentation Platforms in 2026

Mintlify, ReadMe, Stoplight, Redocly and Scalar priced on what a custom domain and white-labeling cost, plus Swagger UI, the open-source core under most.

By DevToolLab Team•

Best WAF and Bot Detection Tools in 2026

Cloudflare, DataDome, HUMAN Security and Akamai compared on real pricing and behavior, plus CrowdSec, the open-source WAF that costs nothing to self-host.

By DevToolLab Team•