AI VRAM Calculator
Most VRAM calculators stop at the model's file size. Running one also needs the KV cache, compute buffers, the driver and your desktop — pick a model, quantisation and context length to see every gigabyte, then which of 33 graphics cards run it, how fast, and what they cost in India.
- Model weights4.5 GB
- 8.03B parameters × 4.85 bits
- KV cache1.0 GB
- 8K context at F16
- Compute buffers0.6 GB
- Scratch space for a 512-token batch
- Driver & runtime0.4 GB
- CUDA / ROCm / Vulkan context
- Desktop & display0.6 GB
- Windows, browser, monitors
- Safety headroom0.7 GB
- 10% for fragmentation and spikes
Enough to load the 4.5 GB model file with the OS and a browser open, and to fall back to offloading if you move up a size. Dual-channel DDR5 matters more than capacity once layers spill out of VRAM.
When the whole model fits in VRAM the processor barely matters — any modern 6-core keeps up. It only matters once layers are offloaded: then a DDR5 platform (AM5 or LGA1851) with fast memory beats extra cores, because generation is limited by memory bandwidth, not compute.
Cards that run it
Cheapest first, at current tracked prices- AMD Radeon RX 6600Fits comfortablyspeed n/a₹19,0138 GB GDDR6
- Intel Arc A580Fits comfortablyspeed n/a₹19,3058 GB GDDR6
- NVIDIA GeForce RTX 3050Fits comfortablyspeed n/a₹21,4508 GB GDDR6
- Intel Arc B570Fits comfortablyspeed n/a₹22,42510 GB GDDR6
- AMD Radeon RX 9060Fits comfortablyspeed n/a₹23,7608 GB GDDR6
- AMD Radeon RX 7600Fits comfortablyspeed n/a₹24,8638 GB GDDR6
AMD: ROCm on Linux, Vulkan or ROCm on Windows. Works in Ollama and LM Studio; expect somewhat lower speeds than an NVIDIA card with the same bandwidth.
INTEL: Vulkan or SYCL backends. Works in LM Studio and llama.cpp; the least mature of the three and the slowest per GB/s.
NVIDIA: CUDA — works out of the box in Ollama, LM Studio and llama.cpp.
Common questions
- Why is VRAM usage higher than the model's file size?
- The file is only the weights. Running it also needs the KV cache — the model's working memory of the conversation, which grows with every token of context — plus compute buffers for the maths in flight, the graphics driver's own context, and whatever your desktop and browser already hold on the card. Together those routinely add 2-3GB to a small model, and far more at long context.
- What is the KV cache and why does context length matter so much?
- For every token in the conversation, each layer of the model stores a key and a value so it does not have to recompute them. That store grows linearly with context: Llama 3.1 8B needs about 128KB per token, so 1GB at 8K tokens and 16GB at 128K — more than three times the model itself. Models with fewer KV heads (grouped-query attention) or sliding-window layers, like Gemma 3, need far less.
- Which quantisation should I use?
- Q4_K_M is the default for good reason: it roughly quarters the memory of the original with a small quality loss. If you have room, Q5_K_M or Q6_K are closer to the original. Below Q4 the loss becomes noticeable, and a smaller model at Q4 usually beats a larger one at Q2 or Q3.
- What happens if the model does not fit in VRAM?
- Ollama and LM Studio will offload some layers to system RAM and run them on the processor. It still works, but those layers run at system-memory speed — roughly a tenth of a modern graphics card's bandwidth — so speed drops sharply even when only a few layers spill over.
- NVIDIA, AMD or Intel for local AI?
- NVIDIA is the path of least resistance: CUDA is supported by every tool out of the box. AMD cards work well in Ollama and LM Studio through ROCm or Vulkan and often offer more VRAM per rupee, but run somewhat slower for the same bandwidth. Intel Arc works through Vulkan or SYCL and is the least mature of the three. For most people, VRAM first, then bandwidth, then brand.
A graphics card needs a machine around it
Big-VRAM cards draw real power and are long. The full builder checks power supply headroom and connectors, card length and slot width against the case, and memory type and capacity against the motherboard.