Model planner

Can I run it?

Pick your hardware, a quantization and a context length. See what fits in memory and the decode ceiling your memory bandwidth allows.

Weights

Most popular, about 4.85 bits per weight.

Context
Memory
12 GB
Published peak
504 GB/s
10 of 22 models run on this machine
  • Llama 3.2 1B
    Fits1.8 of 12 GB
    670tok/s
  • Llama 3.2 3B
    Fits3.2 of 12 GB
    259tok/s
  • Llama 3.1 8B
    Fits6.4 of 12 GB
    104tok/s
  • Llama 3.3 70B
    Too large45.9 of 12 GB
    —
  • Llama 3.1 405B
    Too large252.5 of 12 GB
    —
  • Phi-3.5 mini
    Fits3.6 of 12 GB
    218tok/s
  • Phi-4
    Fits10.7 of 12 GB
    56.6tok/s
  • Gemma 2 9B
    Fits7.2 of 12 GB
    90.0tok/s
  • Gemma 2 27B
    Too large18.7 of 12 GB
    —
  • Mistral 7B
    Fits5.9 of 12 GB
    115tok/s
  • Mistral Small 24B
    Too large16.4 of 12 GB
    —
  • Mixtral 8x7B12.9B active
    Too large31.0 of 12 GB
    —
  • Qwen 2.5 7B
    Fits6.2 of 12 GB
    109tok/s
  • Qwen 2.5 14B
    Fits10.8 of 12 GB
    56.2tok/s
  • Qwen 2.5 32B
    Too large22.2 of 12 GB
    —
  • Qwen 2.5 Coder 32B
    Too large22.2 of 12 GB
    —
  • Qwen 2.5 72B
    Too large47.2 of 12 GB
    —
  • DeepSeek-R1 distill 8B
    Fits6.4 of 12 GB
    104tok/s
  • DeepSeek-R1 distill 32B
    Too large22.2 of 12 GB
    —
  • DeepSeek-R137B active
    Too large414.8 of 12 GB
    —
  • gpt-oss-20b3.6B active
    Too large14.7 of 12 GB
    —
  • gpt-oss-120b5.1B active
    Too large74.6 of 12 GB
    —

The ceiling is memory bandwidth divided by the bytes of weights read per token. It is an upper bound from the memory side; real runtimes land below it. Memory need is the weights plus an estimate for the KV cache and runtime.

Published figures are the ceiling. Yours may differ.

Run the benchmark, then enter your measured read throughput above to plan with what your machine really delivers.