llm.fit

Know what it runs before you buy it.

Choose hardware and a model. The estimate accounts for quantization, KV cache, MoE active parameters, offload, and the multi-GPU interconnect.

Simulation inputs

Paste Hugging Face URL
System RAM
Context
512128k

Estimated decode

32t/s

p10 26p90 39

Fits
weights on GPU 17.3 GBKV 1.5 GBfree 3.7 GBbudget 22.5 GB
Prefill
925 t/s
Time to first token
4.4 s
Memory
18.8 GB
512-token reply
20.2 s

Decode speed vs. context

hover to read
512256k ctxcalibrated →