Qwen3.6 35B A3B Claude 4.7 Opus Reasoning Distilled
Fine-tune of Qwen3.6 35B A3B
Measured axis deltas appear after both rows have board data.
Every measured variant of this model: how much quality each quant keeps, and the VRAM and speed it costs to run.
LB-2026-07.2 | 25/22.5/22.5/22.5/7.5. Where this model and current-lane family runs land vs the frontier anchors.
Swipe horizontally to inspect the full chart →
Complete rows are ordered by Local Intelligence Index; partial rows show their measured axes but are not ranked. The VRAM/Fits columns (8Kcontext) tell you what your card needs. Ranks are within this family's variants.
Variant — quant label plus, where sha-verified, the publisher of the exact weights file benchmarked; rows without a source line predate artifact identity records.
VRAM @8k — model weights, KV cache, and runtime headroom at 8k context.
Fits — the smallest common GPU VRAM tier above that estimate.
Prefill tok/s — prompt-processing speed.
Decode tok/s — generated-token speed after the prompt.
Overall tok/s — completion throughput across the full benchmark.
File size — the benchmarked model artifact on disk.
Runtime — the serving engine and version.
Run — the immutable benchmark receipt. Live rows link to the public submission record until the run is baked into the static site.
Pick the largest quant whose VRAM @8k fits your card.
Swipe horizontally for all variant metrics →
| Rank (this family) | Variant | Local Intelligence IndexLB-2026-07.2 | 25/22.5/22.5/22.5/7.5 | Agentic | Knowledge | Instruction | Coding | Math | VRAM @8k | Fits | Prefill tok/s | Decode tok/s | Overall tok/s | File size | Runtime | Run |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 41.0±2.7 Agentic 6.3 / Knowledge 82.0 / Instruction 57.5 / Coding 28.4 / Math 22.3 | 6.3±5.4 | 82.0±4.2 | 57.5±5.8 | 28.4±7.8 | 22.3±7.2 | 22.1 GB footprint | n/a | 6,098 | 228 | 190.8 | — | llama.cppb9852/fd1a05791 | receipt | |
| — | Q8_0by huihui-ai | no run yet | — | — | — | — | — | 40.6 GB | 48 GB | — | 37.8 GB | — | benchmark it | ||
| — | Q6_Kby huihui-ai | no run yet | — | — | — | — | — | 32 GB | 32 GB | — | 29.2 GB | — | benchmark it | ||
| — | Q2_Kby huihui-ai | no run yet | — | — | — | — | — | 16 GB | 16 GB | — | 13.2 GB | — | benchmark it |
vs base
Fine-tune of Qwen3.6 35B A3B
Measured axis deltas appear after both rows have board data.