GPU/VRAM Calculator for AI

< Tools

Find a GPU for Llama 4, Qwen 3.7, DeepSeek V4, GLM-5.2 and image/video models — or check whether your card has enough VRAM. Accounts for quantization, context length, batch, and light fine-tuning. Linked with the AI catalog and selector.

Context8.192k
2.048k8.192k32.768k131.072k1.048576M
VRAM needed
23 GB
Weights
17.6 GB
KV-cache
2.1 GB
Activations
0.8 GB
Overhead
0.8 GB

Top picks

Best fit
RTX 5090
32 GB · ~$2 000 · Headroom 28%
Most headroom
Mac Studio M2 Ultra 192GB
192 GB · ~$7 000 · Headroom 88%

Compatible GPUs

  • RTX 4090
    24 GB · Ada Lovelace · ~$1 600
    Tight4%
  • RTX 3090
    24 GB · Ampere · ~$800
    Tight4%
  • RTX 3090 Ti
    24 GB · Ampere · ~$900
    Tight4%
  • RTX 5090
    32 GB · Blackwell · ~$2 000
    Fits28%
  • MacBook Pro M4 36GB
    36 GB · Apple Silicon · ~$2 200
    Fits36%
  • A100 40GB
    40 GB · Ampere (Pro) · ~$10 000
    Fits43%
  • L40S
    48 GB · Ada Lovelace (Pro) · ~$7 000
    Fits52%
  • RTX A6000
    48 GB · Ampere (Pro) · ~$4 500
    Fits52%
  • RTX 6000 Ada
    48 GB · Ada Lovelace (Pro) · ~$6 800
    Fits52%
  • MacBook Pro M4 Pro 48GB
    48 GB · Apple Silicon · ~$3 000
    Fits52%
  • A100 80GB
    80 GB · Ampere (Pro) · ~$15 000
    Fits71%
  • H100 80GB
    80 GB · Hopper (Pro) · ~$30 000
    Fits71%
  • MacBook Pro M4 Max 128GB
    128 GB · Apple Silicon · ~$5 000
    Fits82%
  • MacBook Pro M3 Max 128GB
    128 GB · Apple Silicon · ~$4 500
    Fits82%
  • H200 141GB
    141 GB · Hopper (Pro) · ~$35 000
    Fits84%
  • Mac Studio M4 Ultra 192GB
    192 GB · Apple Silicon · ~$8 000
    Fits88%
  • Mac Studio M2 Ultra 192GB
    192 GB · Apple Silicon · ~$7 000
    Fits88%

Estimate ±15–20%. Real usage depends on runtime (llama.cpp, vLLM, Ollama) and drivers. Data verified: 2026-07-23.

Popular model × quant × GPU combos

Indexed starter examples. Numbers come from the calculator formula (inference, batch 1).

GPU/model dataset verified: 2026-07-23

ModelQuantContextVRAMMin GPU
Llama 3.1 8BGGUF Q48.192k7.4 GBRTX 3080 10GB (10 GB)Open
Mistral 7BGGUF Q48.192k6.8 GBRTX 4060 (8 GB)Open
Llama 3.1 70BINT4 / GPTQ / AWQ4.096k41.4 GBA100 80GB (80 GB)Open
Llama 4 MaverickGGUF Q48.192k64.5 GBA100 80GB (80 GB)Open
Qwen 3.7GGUF Q48.192k23 GBRTX 5090 (32 GB)Open
GLM-5.2INT4 / GPTQ / AWQ8.192k59.1 GBA100 80GB (80 GB)Open
FLUX.1 DevFP16 (full)2.048k27.2 GBRTX 5090 (32 GB)Open

VRAM & local LLM FAQ

How much VRAM do I need for a 70B model?
In FP16 a 70B model needs roughly 140 GB just for weights. With INT4 / AWQ / GPTQ you can often land near 35–50 GB at a moderate context (4k–8k), so cards like A100 40GB (tight) or 48–80 GB class GPUs are realistic. Use the calculator with context length enabled — KV-cache grows with tokens.
Q4 vs INT4 — what is the difference?
Both target ~4-bit weights. INT4 usually means GPTQ/AWQ in CUDA stacks (vLLM, transformers). GGUF Q4 is the llama.cpp / Ollama family format, great for consumer GPUs and Apple Silicon. Quality and exact VRAM differ by runtime; Q4_K variants often need slightly more than a raw 0.5 byte/param estimate.
Is 8 GB of VRAM enough for local LLMs?
Yes for smaller models: 7–8B in GGUF Q4 at 2k–8k context typically fits on 8 GB with limited headroom. 13–14B Q4 may be tight. 70B will not fit on 8 GB without heavy offloading. Image models like FLUX.1 Dev usually want 12–24 GB in FP16.
Why does longer context need more VRAM?
KV-cache scales with sequence length × layers × KV heads. Doubling context roughly doubles the KV portion even if weights stay the same. That is why 8B Q4 can jump from ~5–6 GB at 2k to noticeably more at 32k–128k.
How accurate is this calculator?
Treat results as ±15–20%. Real usage depends on runtime (llama.cpp, vLLM, Ollama), CUDA graphs, KV quantization, and driver overhead. Prefer a card with comfortable headroom rather than a perfect tight fit.

Related