GPU Custom RTX 3060 (12GB) — 360 GB/s, 51.2 TFLOPS RTX 3060 Ti (8GB) — 448 GB/s, 64.8 TFLOPS RTX 3070 (8GB) — 448 GB/s, 81.2 TFLOPS RTX 3070 Ti (8GB) — 608 GB/s, 87 TFLOPS RTX 3080 (10GB) — 760.3 GB/s, 119.1 TFLOPS RTX 3080 (12GB) — 912 GB/s, 122.6 TFLOPS RTX 3080 Ti (12GB) — 912.4 GB/s, 136.4 TFLOPS RTX 3090 (24GB) — 936.2 GB/s, 142.8 TFLOPS RTX 3090 Ti (24GB) — 1008 GB/s, 160 TFLOPS RTX 4060 (8GB) — 272 GB/s, 60.4 TFLOPS RTX 4060 Ti (8GB) — 288 GB/s, 88.2 TFLOPS RTX 4060 Ti (16GB) — 288 GB/s, 88.2 TFLOPS RTX 4070 (12GB) — 504 GB/s, 116.6 TFLOPS RTX 4070 Ti (12GB) — 504.2 GB/s, 160.4 TFLOPS RTX 4070 Ti SUPER (16GB) — 672.3 GB/s, 176.4 TFLOPS RTX 4070 SUPER (12GB) — 504.2 GB/s, 141.9 TFLOPS RTX 4080 (16GB) — 716.8 GB/s, 195 TFLOPS RTX 4080 SUPER (16GB) — 736.3 GB/s, 208.9 TFLOPS RTX 4090 (24GB) — 1008 GB/s, 165 TFLOPS RTX 5060 (8GB) — 448 GB/s, 76.7 TFLOPS RTX 5060 Ti (8GB) — 448 GB/s, 94.8 TFLOPS RTX 5060 Ti (16GB) — 448 GB/s, 94.8 TFLOPS RTX 5070 (12GB) — 672 GB/s, 123.5 TFLOPS RTX 5070 Ti (16GB) — 896 GB/s, 175.8 TFLOPS RTX 5080 (16GB) — 960 GB/s, 225.1 TFLOPS RTX 5090 (32GB) — 1792 GB/s, 419.2 TFLOPS RTX 2000 Ada Generation (16GB) — 224 GB/s, 48 TFLOPS RTX 4000 Ada Generation (20GB) — 360 GB/s, 106.9 TFLOPS RTX 4500 Ada Generation (24GB) — 432 GB/s, 158.4 TFLOPS RTX 5000 Ada Generation (32GB) — 576 GB/s, 261.1 TFLOPS RTX 6000 Ada Generation (48GB) — 960 GB/s, 364.4 TFLOPS RTX A4000 (16GB) — 448 GB/s, 76.8 TFLOPS RTX A5000 (24GB) — 768 GB/s, 111.1 TFLOPS RTX A6000 (48GB) — 768 GB/s, 154.8 TFLOPS RTX PRO 4000 Blackwell (24GB) — 672 GB/s, 160 TFLOPS RTX PRO 4000 Blackwell SFF (24GB) — 432 GB/s, 96.2 TFLOPS RTX PRO 4500 Blackwell (32GB) — 896 GB/s, 203 TFLOPS RTX PRO 5000 Blackwell (48GB) — 1344 GB/s, 267.8 TFLOPS RTX PRO 6000 Blackwell (96GB) — 1792 GB/s, 504 TFLOPS L40 (48GB) — 864 GB/s, 181 TFLOPS L40S (48GB) — 864 GB/s, 362.05 TFLOPS L4 (24GB) — 300 GB/s, 121 TFLOPS A10G (24GB) — 600 GB/s, 70 TFLOPS T4 (16GB) — 320 GB/s, 65 TFLOPS A2 (16GB) — 200 GB/s, 18 TFLOPS A16 (64GB) — 800 GB/s, 72 TFLOPS A30 (24GB) — 933 GB/s, 165 TFLOPS A40 (48GB) — 696 GB/s, 149.7 TFLOPS A100 (40GB) — 1555 GB/s, 312 TFLOPS A100 (80GB) — 2039 GB/s, 312 TFLOPS A800 (40GB) — 1555.2 GB/s, 78 TFLOPS A800 (80GB) — 1935 GB/s, 312 TFLOPS H100 (80GB) — 3350 GB/s, 989.5 TFLOPS H100 NVL (94GB) — 3938 GB/s, 835.5 TFLOPS H200 (141GB) — 4800 GB/s, 989.5 TFLOPS H800 (80GB) — 2000 GB/s, 756.5 TFLOPS B100 (192GB) — 8000 GB/s, 1750 TFLOPS B200 (192GB) — 8000 GB/s, 2250 TFLOPS Blackwell B300 (288GB) — 8000 GB/s, 2250 TFLOPS Rubin R100 (288GB) — 22000 GB/s, 4000 TFLOPS GB10 Grace Blackwell (128GB) — DGX Spark — 273.2 GB/s, 150 TFLOPS RX 6600 (8GB) — 224 GB/s, 18 TFLOPS RX 6600 XT (8GB) — 256 GB/s, 21 TFLOPS RX 6650 XT (8GB) — 280 GB/s, 22 TFLOPS RX 6700 (10GB) — 320 GB/s, 22.6 TFLOPS RX 6700 XT (12GB) — 384 GB/s, 26 TFLOPS RX 6750 XT (12GB) — 432 GB/s, 27 TFLOPS RX 6800 (16GB) — 512 GB/s, 32 TFLOPS RX 6800 XT (16GB) — 512 GB/s, 41 TFLOPS RX 6900 XT (16GB) — 512 GB/s, 46 TFLOPS RX 6950 XT (16GB) — 576 GB/s, 47 TFLOPS RX 7600 (8GB) — 288 GB/s, 43 TFLOPS RX 7600 XT (16GB) — 288 GB/s, 45 TFLOPS RX 7700 XT (12GB) — 432 GB/s, 70 TFLOPS RX 7800 XT (16GB) — 624 GB/s, 74 TFLOPS RX 7900 GRE (16GB) — 576 GB/s, 92 TFLOPS RX 7900 XT (20GB) — 800 GB/s, 103 TFLOPS RX 7900 XTX (24GB) — 960 GB/s, 123 TFLOPS RX 9070 (16GB) — 640 GB/s, 145 TFLOPS RX 9070 XT (16GB) — 640 GB/s, 194 TFLOPS Radeon PRO W6600 (8GB) — 224 GB/s, 20 TFLOPS Radeon PRO W6800 (32GB) — 512 GB/s, 35 TFLOPS Radeon PRO W7800 (32GB) — 576 GB/s, 90 TFLOPS Radeon PRO W7900 (48GB) — 864 GB/s, 120 TFLOPS Instinct MI210 (64GB) — 1638 GB/s, 181 TFLOPS Instinct MI250 (128GB) — 3276 GB/s, 362 TFLOPS Instinct MI250X (128GB) — 3276 GB/s, 383 TFLOPS Instinct MI300A (128GB) — 5300 GB/s, 980.6 TFLOPS Instinct MI300X (192GB) — 5300 GB/s, 1300 TFLOPS Instinct MI325X (256GB) — 6000 GB/s, 1300 TFLOPS Ryzen AI MAX+ 395 (128GB system, 96GB GPU-accessible) — 256 GB/s, 59.392 TFLOPS Ryzen AI 9 HX 370 (32GB system, 24GB GPU-accessible) — 89.6 GB/s, 11.88 TFLOPS Ryzen AI 9 HX 370 (64GB system, 48GB GPU-accessible) — 102.4 GB/s, 11.88 TFLOPS Ryzen AI 9 HX 370 (96GB system, 72GB GPU-accessible) — 128 GB/s, 11.88 TFLOPS M2 Pro (16GB) — 200 GB/s, 11.36 TFLOPS M2 Max (32GB) — 400 GB/s, 27.2 TFLOPS M2 Max (64GB) — 400 GB/s, 27.2 TFLOPS M2 Max (96GB) — 400 GB/s, 27.2 TFLOPS M2 Ultra (64GB) — 800 GB/s, 54.4 TFLOPS M2 Ultra (128GB) — 800 GB/s, 54.4 TFLOPS M2 Ultra (192GB) — 800 GB/s, 54.4 TFLOPS M3 Pro (18GB) — 150 GB/s, 12.78 TFLOPS M3 Pro (36GB) — 150 GB/s, 12.78 TFLOPS M3 Max (36GB) — 300 GB/s, 21.3 TFLOPS M3 Max (48GB) — 400 GB/s, 32.8 TFLOPS M3 Max (64GB) — 400 GB/s, 32.8 TFLOPS M3 Max (96GB) — 400 GB/s, 32.8 TFLOPS M3 Max (128GB) — 400 GB/s, 32.8 TFLOPS M3 Ultra (256GB) — 819.2 GB/s, 56.8 TFLOPS M3 Ultra (512GB) — 819.2 GB/s, 56.8 TFLOPS M4 (16GB) — 120 GB/s, 8.5 TFLOPS M4 (24GB) — 120 GB/s, 8.5 TFLOPS M4 (32GB) — 120 GB/s, 8.5 TFLOPS M4 Pro (32GB) — 273 GB/s, 18.4 TFLOPS M4 Pro (48GB) — 273 GB/s, 18.4 TFLOPS M4 Pro (64GB) — 273 GB/s, 18.4 TFLOPS M4 Max (64GB) — 546 GB/s, 34.08 TFLOPS M4 Max (96GB) — 546 GB/s, 34.08 TFLOPS M4 Max (128GB) — 546 GB/s, 34.08 TFLOPS M5 (16GB) — 153 GB/s, 8.3 TFLOPS M5 (24GB) — 153 GB/s, 8.3 TFLOPS M5 (32GB) — 153 GB/s, 8.3 TFLOPS M5 Pro (24GB) — 307 GB/s, 16.6 TFLOPS M5 Pro (48GB) — 307 GB/s, 16.6 TFLOPS M5 Pro (64GB) — 307 GB/s, 16.6 TFLOPS M5 Max (48GB) — 614 GB/s, 33.2 TFLOPS M5 Max (64GB) — 614 GB/s, 33.2 TFLOPS M5 Max (96GB) — 614 GB/s, 33.2 TFLOPS M5 Max (128GB) — 614 GB/s, 33.2 TFLOPS
VRAM/GPU (GB)
Num GPUs
Bandwidth (GB/s)
FP16 TFLOPS
Overhead reserve (%)
Serving engine vLLM Ollama llama.cpp
vLLM: set --gpu-memory-utilization per instance so the sum across instances stays under 1.0 minus overhead.
Formulas — Weights (GB) = params (B) × bytes/param KV cache/token (bytes) = 2 × layers × kv_heads × head_dim × bytes/param(KV) KV cache/user (GB) = KV cache/token × context tokens ÷ 1e9 Total KV cache (GB) = KV cache/user × concurrent users
Notes — Architecture presets (layers / kv_heads / head_dim) are approximate — confirm exact values from the model's config.json on Hugging Face before final capacity planning. For MoE models, weights sizing uses total parameters (every expert must be resident in VRAM even though only some activate per token). Models marked "too large for 1 GPU" are multi-hundred-billion to multi-trillion parameter MoE models included so you can see exactly how far over budget they'd be on a single card. vLLM's PagedAttention allocates KV cache in blocks rather than pre-reserving full context length per user, so real vLLM usage often runs a bit lower than this estimate; Ollama pre-allocates per OLLAMA_NUM_PARALLEL slot at num_ctx size, so it tracks this estimate more closely.
Performance estimate — Decode/aggregate throughput and time-to-first-token use a simple memory-bandwidth / compute "roofline" model (decode is bandwidth-bound, prefill is compute-bound), based on each GPU's peak theoretical bandwidth and dense FP16 TFLOPS. This is not benchmark-calibrated — real throughput depends heavily on the serving engine's efficiency (MFU), batching strategy, quantization kernel quality, and PCIe/NVLink interconnect for multi-GPU setups, and is typically 30–60% below this theoretical ceiling in practice. Treat these numbers as an order-of-magnitude sanity check, not a substitute for benchmarking on real hardware.
Fine-tuning estimate — Covers static memory only — weights, gradients, and optimizer state. Full fine-tuning uses a standard mixed-precision AdamW estimate (~16 bytes/param: bf16 weight + bf16 grad + fp32 master weight + fp32 Adam m + fp32 Adam v). LoRA/QLoRA freeze the base model (no gradients/optimizer on those weights) and estimate the trainable adapter at ~12 bytes per trainable parameter, based on the "Trainable %" you set — this is a rough proxy for LoRA rank/target-module choice, not an exact adapter-size calculation. Not included: activation memory, which scales with batch size and sequence length and can add a substantial amount on top for full fine-tuning without gradient checkpointing — budget extra headroom rather than treating this number as a hard ceiling.