LocalAI Cost — Buy the card or rent the tokens?
Model
Total params (B)
Active params (B)
KV per 1K tokens (GB)
Quant
BF16 / FP16 (2.0)
FP8 (1.0)
Q8_0 (~1.06)
Q6_K (~0.75)
Q5_K_M (~0.69)
Q4_K_M (~0.6)
MXFP4 (~0.55)
Q3_K_M (~0.47)
Q2_K (~0.4)
Your usage
Usage hours / week
Electricity (USD/kWh)
Working context
128K
256K
512K
1M
System prompt (K tokens)
Hosted out (USD / M tok)
Measured t/s (override, blank = estimate)
t/s override applies to every card row. Working context and hosted out are prefilled from the model preset — edit for your real values.
Card
VRAM (used / total)
Bandwidth
Verdict
t/s
First token
Break-even
Price
Accuracy