LLM Inference: VRAM & Performance Calculator

Pick a model, a quantization and a device — see whether it fits and how fast it runs.

Precision of the model weights. Lower uses less VRAM but costs quality.

KV cache precision. Dominates VRAM at long context.

Hardware

Select your GPU or configure a custom device.

Number of devices1

Tensor-parallel replicas. Comms overhead is included.

12481632

Lets the model exceed VRAM — at host-bandwidth speed.

Workload

Batch size1

Sequences processed per step. Raises throughput, costs KV cache.

141664128
Sequence length1,024

Tokens per sequence (prompt + generation). Drives KV cache.

Concurrent users1

Simultaneous requests. Multiplies KV cache, splits per-user speed.

141664128
0%of VRAM
Comfortable

0 GB

of 0 GB usable

Generation speed
Per-token latency
Time to first token
Total throughput
Bottleneck

Memory allocation