GPU/VRAM Calculator for AI
< ToolsFind a GPU for Llama 4, Qwen 3.7, DeepSeek V4, GLM-5.2 and image/video models — or check whether your card has enough VRAM. Accounts for quantization, context length, batch, and light fine-tuning. Linked with the AI catalog and selector.
Context8.192k
2.048k8.192k32.768k131.072k1.048576M
VRAM needed
7.4 GB
Weights
4.4 GB
KV-cache
1.1 GB
Activations
0.6 GB
Overhead
0.8 GB
Top picks
Best fit
RTX 3080 10GB
10 GB · ~$400 · Headroom 26%
Cheapest
RTX 3060 12GB
12 GB · ~$250 · Headroom 38%
Most headroom
Mac Studio M2 Ultra 192GB
192 GB · ~$7 000 · Headroom 96%
Compatible GPUs
- RTX 5060 Ti 8GB8 GB · Blackwell · ~$380Tight7%
- RTX 4060 Ti 8GB8 GB · Ada Lovelace · ~$400Tight7%
- RTX 40608 GB · Ada Lovelace · ~$300Tight7%
- RTX 3070 Ti8 GB · Ampere · ~$350Tight7%
- RTX 3080 10GB10 GB · Ampere · ~$400Fits26%
- RTX 507012 GB · Blackwell · ~$550Fits38%
- RTX 4070 Super12 GB · Ada Lovelace · ~$600Fits38%
- RTX 407012 GB · Ada Lovelace · ~$500Fits38%
- RTX 3080 Ti12 GB · Ampere · ~$500Fits38%
- RTX 3080 12GB12 GB · Ampere · ~$450Fits38%
- RTX 3060 12GB12 GB · Ampere · ~$250Fits38%
- RTX 508016 GB · Blackwell · ~$1 000Fits54%
- RTX 5070 Ti16 GB · Blackwell · ~$750Fits54%
- RTX 5060 Ti 16GB16 GB · Blackwell · ~$430Fits54%
- RTX 4080 Super16 GB · Ada Lovelace · ~$1 000Fits54%
- RTX 4070 Ti Super16 GB · Ada Lovelace · ~$800Fits54%
- RTX 4060 Ti 16GB16 GB · Ada Lovelace · ~$450Fits54%
- RTX 409024 GB · Ada Lovelace · ~$1 600Fits69%
- RTX 309024 GB · Ampere · ~$800Fits69%
- RTX 3090 Ti24 GB · Ampere · ~$900Fits69%
- RTX 509032 GB · Blackwell · ~$2 000Fits77%
- MacBook Pro M4 36GB36 GB · Apple Silicon · ~$2 200Fits79%
- A100 40GB40 GB · Ampere (Pro) · ~$10 000Fits82%
- L40S48 GB · Ada Lovelace (Pro) · ~$7 000Fits85%
- RTX A600048 GB · Ampere (Pro) · ~$4 500Fits85%
- RTX 6000 Ada48 GB · Ada Lovelace (Pro) · ~$6 800Fits85%
- MacBook Pro M4 Pro 48GB48 GB · Apple Silicon · ~$3 000Fits85%
- A100 80GB80 GB · Ampere (Pro) · ~$15 000Fits91%
- H100 80GB80 GB · Hopper (Pro) · ~$30 000Fits91%
- MacBook Pro M4 Max 128GB128 GB · Apple Silicon · ~$5 000Fits94%
- MacBook Pro M3 Max 128GB128 GB · Apple Silicon · ~$4 500Fits94%
- H200 141GB141 GB · Hopper (Pro) · ~$35 000Fits95%
- Mac Studio M4 Ultra 192GB192 GB · Apple Silicon · ~$8 000Fits96%
- Mac Studio M2 Ultra 192GB192 GB · Apple Silicon · ~$7 000Fits96%
Estimate ±15–20%. Real usage depends on runtime (llama.cpp, vLLM, Ollama) and drivers. Data verified: 2026-07-23.
Popular model × quant × GPU combos
Indexed starter examples. Numbers come from the calculator formula (inference, batch 1).
GPU/model dataset verified: 2026-07-23
| Model | Quant | Context | VRAM | Min GPU | |
|---|---|---|---|---|---|
| Llama 3.1 8B | GGUF Q4 | 8.192k | 7.4 GB | RTX 3080 10GB (10 GB) | Open |
| Mistral 7B | GGUF Q4 | 8.192k | 6.8 GB | RTX 4060 (8 GB) | Open |
| Llama 3.1 70B | INT4 / GPTQ / AWQ | 4.096k | 41.4 GB | A100 80GB (80 GB) | Open |
| Llama 4 Maverick | GGUF Q4 | 8.192k | 64.5 GB | A100 80GB (80 GB) | Open |
| Qwen 3.7 | GGUF Q4 | 8.192k | 23 GB | RTX 5090 (32 GB) | Open |
| GLM-5.2 | INT4 / GPTQ / AWQ | 8.192k | 59.1 GB | A100 80GB (80 GB) | Open |
| FLUX.1 Dev | FP16 (full) | 2.048k | 27.2 GB | RTX 5090 (32 GB) | Open |
VRAM & local LLM FAQ
- How much VRAM do I need for a 70B model?
- In FP16 a 70B model needs roughly 140 GB just for weights. With INT4 / AWQ / GPTQ you can often land near 35–50 GB at a moderate context (4k–8k), so cards like A100 40GB (tight) or 48–80 GB class GPUs are realistic. Use the calculator with context length enabled — KV-cache grows with tokens.
- Q4 vs INT4 — what is the difference?
- Both target ~4-bit weights. INT4 usually means GPTQ/AWQ in CUDA stacks (vLLM, transformers). GGUF Q4 is the llama.cpp / Ollama family format, great for consumer GPUs and Apple Silicon. Quality and exact VRAM differ by runtime; Q4_K variants often need slightly more than a raw 0.5 byte/param estimate.
- Is 8 GB of VRAM enough for local LLMs?
- Yes for smaller models: 7–8B in GGUF Q4 at 2k–8k context typically fits on 8 GB with limited headroom. 13–14B Q4 may be tight. 70B will not fit on 8 GB without heavy offloading. Image models like FLUX.1 Dev usually want 12–24 GB in FP16.
- Why does longer context need more VRAM?
- KV-cache scales with sequence length × layers × KV heads. Doubling context roughly doubles the KV portion even if weights stay the same. That is why 8B Q4 can jump from ~5–6 GB at 2k to noticeably more at 32k–128k.
- How accurate is this calculator?
- Treat results as ±15–20%. Real usage depends on runtime (llama.cpp, vLLM, Ollama), CUDA graphs, KV quantization, and driver overhead. Prefer a card with comfortable headroom rather than a perfect tight fit.