Required for the site to work (login, preferences, security). Always on.
Which AI model can your hardware run?
Pick a model, a quantization level and how many people will use it at once — get an honest memory estimate, which GPUs or Macs actually fit, and whether buying hardware beats paying for API tokens. Nothing here is sent to a server: it all runs in your browser.
1. Configure
Dense-model assumption — for your own fine-tune or a model not in the list.
2. Memory needed
How we calculate this
Weights ≈ parameters × bits per weight ÷ 8 × 1.1 (a 10% allowance for runtime buffers and activations).
"KV width" is the model's attention output size after grouped-query reduction — published in some model configs, approximated from comparable architectures for others (see sources below).
DeepSeek-V3 compresses key and value into one shared latent vector per token (Multi-head Latent Attention) instead of two full-width caches, so its KV cache is far smaller than the generic formula above — we use the compressed dimension from its technical report.
These are estimates, not guarantees. Real usage varies with the inference engine, batching strategy and prompt shape.
3. What fits
Memory fit only — we never invent tokens-per-second numbers we can't measure.
| Hardware | Memory | Fit |
|---|---|---|
| RTX 4060 Ti 16GB | 16 GB | Fits |
| RTX 4090 24GB | 24 GB | Fits |
| RTX 5090 32GB | 32 GB | Fits |
| 2× RTX 4090 (48GB) | 48 GB | Fits |
| RTX 6000-class 48GB | 48 GB | Fits |
| RTX 6000-class 96GB | 96 GB | Fits |
| Apple M-series 32GB | 32 GB | Fits |
| Apple M-series 64GB | 64 GB | Fits |
| Apple M-series 128GB | 128 GB | Fits |
| NVIDIA DGX Spark-class 128GB | 128 GB | Fits |
"Fits" leaves headroom for the OS and other apps. "Tight" means it's within 15% over that headroom — a smaller context or heavier quantization would close the gap.
4. Buy hardware or pay per token?
Edit any number below — these are starting points, not our prices.
Amortized over 36 months
Assumptions
- Memory estimates assume the whole model and its KV cache sit in VRAM or unified memory at once — no offloading to system RAM.
- The 1.1× overhead factor is a rough allowance for runtime buffers and activation memory, not a measured figure for any specific inference engine.
- KV-cache width is confirmed from the model's own configuration for some entries and approximated from comparable architectures for others — see sources.
- Quantization below Q8 changes model behavior, not just file size — always check quality for your own use case.
- Hardware fit compares required memory to a fixed usable-memory assumption per device; actual headroom depends on your OS, drivers and what else is running.
- The cost comparison ignores maintenance, depreciation beyond 36 months, and the value of your own time.
Sources
Parameter counts are checked against each model's official card, technical report or Hugging Face configuration:
- Llama 3.1 8B — Meta — Llama 3 model card
- Llama 3.1 70B — Meta — Llama 3.1 70B (Hugging Face config)
- Llama 3.3 70B — Meta — Llama 3.3 70B (same architecture as 3.1 70B)
- Qwen3 8B — Alibaba — Qwen3-8B model card
- Qwen3 32B — Alibaba — Qwen3-32B model card
- Qwen3 235B-A22B — Alibaba — Qwen3-235B-A22B model card
- Gemma 3 12B — Google — Gemma 3 Technical Report (Table 1)
- Gemma 3 27B — Google — Gemma 3 Technical Report (Table 1)
- Mistral Small 24B — Mistral AI — Mistral Small 3.1/3.2 (24B) model card
- gpt-oss-20b — OpenAI — gpt-oss model card (Table 1)
- gpt-oss-120b — OpenAI — gpt-oss model card (Table 1)
- DeepSeek-V3 — DeepSeek — DeepSeek-V3 technical report + model card
- GLM-4.5-Air — Z.ai — GLM-4.5-Air model card
- GLM-4.5 — Z.ai — GLM-4.5 model card
- Llama 3.1 405B — Meta — Llama 3.1 405B (Hugging Face config)
Need this running in production, not just estimated?
Altovar deploys sovereign, local AI for European businesses — on your own hardware or in an EU data center, with no data leaving your infrastructure. See how Altovar deploys local AI →
New to running models locally? 3 repos every SME should know for local AI →