Know what it runs before you buy it.
Choose hardware and a model. The estimate accounts for quantization, KV cache, MoE active parameters, offload, and the multi-GPU interconnect.
Simulation inputs
Paste Hugging Face URL
System RAM
Context
512128k
Estimated decode
32t/s
p10 26p90 39
weights on GPU 17.3 GBKV 1.5 GBfree 3.7 GBbudget 22.5 GB
- Prefill
- 925 t/s
- Time to first token
- 4.4 s
- Memory
- 18.8 GB
- 512-token reply
- 20.2 s