DeepSeek V3.2
Every measured variant of this model: how much quality each quant keeps, and the VRAM and speed it costs to run.
VRAM footprint vs Local Intelligence Index
LB-2026-07.2 | 25/22.5/22.5/22.5/7.5
This chart appears after the model's first measured run lands. Use the quant ladder below for file size and VRAM requirements.
Variant profiles
Complete rows are ordered by Local Intelligence Index; partial rows show their measured axes but are not ranked. The VRAM/Fits columns (8Kcontext) tell you what your card needs. Ranks are within this family's variants.
Column guide
Variant — quant label plus, where sha-verified, the publisher of the exact weights file benchmarked; rows without a source line predate artifact identity records.
VRAM @8k — model weights, KV cache, and runtime headroom at 8k context.
Fits — the smallest common GPU VRAM tier above that estimate.
Prefill tok/s — prompt-processing speed.
Decode tok/s — generated-token speed after the prompt.
Overall tok/s — completion throughput across the full benchmark.
File size — the benchmarked model artifact on disk.
Runtime — the serving engine and version.
Run — the immutable benchmark receipt. Live rows link to the public submission record until the run is baked into the static site.
Pick the largest quant whose VRAM @8k fits your card.
Swipe horizontally for all variant metrics →
| Rank (this family) | Variant | Local Intelligence Indexindex-v3.0 | 40/15/15/10/15/5 | VRAM @8k | Fits | Overall tok/s | File size | Runtime | Run |
|---|---|---|---|---|---|---|---|---|
| — | Q8_0 | no run yet | 747.8 GB | >512 GB | — | 713.3 GB | — | benchmark it |
| — | Q6_K | no run yet | 585.8 GB | >512 GB | — | 551.3 GB | — | benchmark it |
| — | Q5_K_M | no run yet | 510.8 GB | 512 GB | — | 476.2 GB | — | benchmark it |
| — | Q4_K_M | no run yet | 439.9 GB | 512 GB | — | 405.4 GB | — | benchmark it |
| — | Q3_K_M | no run yet | 354.8 GB | 384 GB | — | 320.2 GB | — | benchmark it |
| — | Q2_K | no run yet | 279.9 GB | 384 GB | — | 245.4 GB | — | benchmark it |