FREE TOOL

Which AI model can your hardware run?

Pick a model, a quantization level and how many people will use it at once — get an honest memory estimate, which GPUs or Macs actually fit, and whether buying hardware beats paying for API tokens. Nothing here is sent to a server: it all runs in your browser.

1. Configure

2. Memory needed

Model weights4.10 GB
KV cache1.00 GB
Total required5.10 GB
How we calculate this

Weights ≈ parameters × bits per weight ÷ 8 × 1.1 (a 10% allowance for runtime buffers and activations).

"KV width" is the model's attention output size after grouped-query reduction — published in some model configs, approximated from comparable architectures for others (see sources below).

These are estimates, not guarantees. Real usage varies with the inference engine, batching strategy and prompt shape.

3. What fits

Memory fit only — we never invent tokens-per-second numbers we can't measure.

HardwareMemoryFit
RTX 4060 Ti 16GB 16 GB Fits
RTX 4090 24GB 24 GB Fits
RTX 5090 32GB 32 GB Fits
2× RTX 4090 (48GB) 48 GB Fits
RTX 6000-class 48GB 48 GB Fits
RTX 6000-class 96GB 96 GB Fits
Apple M-series 32GB 32 GB Fits
Apple M-series 64GB 64 GB Fits
Apple M-series 128GB 128 GB Fits
NVIDIA DGX Spark-class 128GB 128 GB Fits

"Fits" leaves headroom for the OS and other apps. "Tight" means it's within 15% over that headroom — a smaller context or heavier quantization would close the gap.

4. Buy hardware or pay per token?

Edit any number below — these are starting points, not our prices.

Amortized over 36 months

Monthly API cost€60.00
Monthly local cost (amortized hardware + electricity)€81.13
Break-evenBeyond the 36-month amortization window at this usage

Assumptions

  • Memory estimates assume the whole model and its KV cache sit in VRAM or unified memory at once — no offloading to system RAM.
  • The 1.1× overhead factor is a rough allowance for runtime buffers and activation memory, not a measured figure for any specific inference engine.
  • KV-cache width is confirmed from the model's own configuration for some entries and approximated from comparable architectures for others — see sources.
  • Quantization below Q8 changes model behavior, not just file size — always check quality for your own use case.
  • Hardware fit compares required memory to a fixed usable-memory assumption per device; actual headroom depends on your OS, drivers and what else is running.
  • The cost comparison ignores maintenance, depreciation beyond 36 months, and the value of your own time.

Sources

Parameter counts are checked against each model's official card, technical report or Hugging Face configuration:

Need this running in production, not just estimated?

Altovar deploys sovereign, local AI for European businesses — on your own hardware or in an EU data center, with no data leaving your infrastructure. See how Altovar deploys local AI →

New to running models locally? 3 repos every SME should know for local AI →