VRAM Capacity Planner

Estimate model-weight and KV-cache footprint per model, and see how many concurrent users fit before the card runs out, for Ollama, Llama CPP or vLLM deployments sharing one GPU.

 

vLLM: set --gpu-memory-utilization per instance so the sum across instances stays under 1.0 minus overhead.

GPU allocation summary

62.1 / 96 GB used

Model weightsKV cacheOverhead reserveFree

Models on this GPU

Presets are approximate architecture values — verify against the model's config.json.

Weights17.6 GB
KV / user2.10 GB
KV total (15 users)31.5 GB
Model total49.1 GB
Decode (1 user)102 tok/s
Aggregate (15 users)1527 tok/s
Time to first token1016 ms
Compute ceiling7875 tok/s

Presets are approximate architecture values — verify against the model's config.json.

Weights1.2 GB
KV / user0.02 GB
KV total (15 users)0.4 GB
Model total1.6 GB
Decode (1 user)1493 tok/s
Aggregate (15 users)22400 tok/s
Time to first token5 ms
Compute ceiling420000 tok/s

Clients & Partners

18+ companies
Agro Farm
Alpha Chemical
Bloom Corp
Bloom Loan
Carson Mountains
Eternal Eye
Glass Factory
Risings Education
Safirify
Nexook
Nutribest
XP3 Data
Agro Farm
Alpha Chemical
Bloom Corp
Bloom Loan
Carson Mountains
Eternal Eye
Glass Factory
Risings Education
Safirify
Nexook
Nutribest
XP3 Data
● trusted by industry leaders