AI VRAM Calculator

Most VRAM calculators stop at the model's file size. Running one also needs the KV cache, compute buffers, the driver and your desktop — pick a model, quantisation and context length to see every gigabyte, then which of 33 graphics cards run it, how fast, and what they cost in India.

The default sweet spot — what Ollama and LM Studio ship.
What Ollama and LM Studio use unless told otherwise.
8K tokens
2K4K8K16K32K64K128K
Roughly ¾ of a word per token — 8K is a long chat, 32K a short book chapter, 128K a whole codebase.
Each parallel chat gets its own KV cache (Ollama's OLLAMA_NUM_PARALLEL).
Minimum
7.1 GB
Loads and runs, nothing to spare
Recommended
7.8 GB
With 10% headroom
Buy a card with
8 GB
Smallest size sold that holds it
Model weights4.5 GB
8.03B parameters × 4.85 bits
KV cache1.0 GB
8K context at F16
Compute buffers0.6 GB
Scratch space for a 512-token batch
Driver & runtime0.4 GB
CUDA / ROCm / Vulkan context
Desktop & display0.6 GB
Windows, browser, monitors
Safety headroom0.7 GB
10% for fragmentation and spikes
System RAM
16 GB

Enough to load the 4.5 GB model file with the OS and a browser open, and to fall back to offloading if you move up a size. Dual-channel DDR5 matters more than capacity once layers spill out of VRAM.

Processor

When the whole model fits in VRAM the processor barely matters — any modern 6-core keeps up. It only matters once layers are offloaded: then a DDR5 platform (AM5 or LGA1851) with fast memory beats extra cores, because generation is limited by memory bandwidth, not compute.

Cards that run it

Cheapest first, at current tracked prices

AMD: ROCm on Linux, Vulkan or ROCm on Windows. Works in Ollama and LM Studio; expect somewhat lower speeds than an NVIDIA card with the same bandwidth.

INTEL: Vulkan or SYCL backends. Works in LM Studio and llama.cpp; the least mature of the three and the slowest per GB/s.

NVIDIA: CUDA — works out of the box in Ollama, LM Studio and llama.cpp.

Build a PC around it
Estimates for llama.cpp, Ollama and LM Studio with GGUF models and flash attention on — expect real usage within about 10%. Speeds are for generating a reply, in one chat, with the conversation half full; reading a long prompt is much faster. They come from each card's memory bandwidth, not a benchmark, so treat them as a way to compare cards rather than a promise.

Common questions

Why is VRAM usage higher than the model's file size?
The file is only the weights. Running it also needs the KV cache — the model's working memory of the conversation, which grows with every token of context — plus compute buffers for the maths in flight, the graphics driver's own context, and whatever your desktop and browser already hold on the card. Together those routinely add 2-3GB to a small model, and far more at long context.
What is the KV cache and why does context length matter so much?
For every token in the conversation, each layer of the model stores a key and a value so it does not have to recompute them. That store grows linearly with context: Llama 3.1 8B needs about 128KB per token, so 1GB at 8K tokens and 16GB at 128K — more than three times the model itself. Models with fewer KV heads (grouped-query attention) or sliding-window layers, like Gemma 3, need far less.
Which quantisation should I use?
Q4_K_M is the default for good reason: it roughly quarters the memory of the original with a small quality loss. If you have room, Q5_K_M or Q6_K are closer to the original. Below Q4 the loss becomes noticeable, and a smaller model at Q4 usually beats a larger one at Q2 or Q3.
What happens if the model does not fit in VRAM?
Ollama and LM Studio will offload some layers to system RAM and run them on the processor. It still works, but those layers run at system-memory speed — roughly a tenth of a modern graphics card's bandwidth — so speed drops sharply even when only a few layers spill over.
NVIDIA, AMD or Intel for local AI?
NVIDIA is the path of least resistance: CUDA is supported by every tool out of the box. AMD cards work well in Ollama and LM Studio through ROCm or Vulkan and often offer more VRAM per rupee, but run somewhat slower for the same bandwidth. Intel Arc works through Vulkan or SYCL and is the least mature of the three. For most people, VRAM first, then bandwidth, then brand.

A graphics card needs a machine around it

Big-VRAM cards draw real power and are long. The full builder checks power supply headroom and connectors, card length and slot width against the case, and memory type and capacity against the motherboard.