How much VRAM does a model need? Estimate memory requirements at every quantization level — FP32 down to 1.58-bit — and check hardware fit instantly.
Runs entirely in your browser. Nothing is sent anywhere.
Enter model size or select a known architecture. Overhead adds ~10–20% to the raw weight footprint for KV cache, activations, and framework buffers.
Quick select
| Format | Bits/param | Raw (GB) | With overhead | Hardware fit |
|---|
| GPU / Setup | VRAM | Recommended max load |
|---|---|---|
| RTX 3060 | 12 GB | ~10 GB |
| RTX 4090 | 24 GB | ~20 GB |
| Kaggle T4 ×1 | 16 GB | ~13 GB |
| Kaggle T4 ×2 | 32 GB | ~26 GB |
| A100 40 GB | 40 GB | ~34 GB |
| A100 80 GB | 80 GB | ~68 GB |
| CPU only (RAM) | varies | depends on system RAM |