Skip to content

mixed / LB-2026-07.2

Global comparison

Every ranked model is scored on the same frozen suite across Agentic, Knowledge, Instruction, Coding, Math. This is the initial measured ladder — more models land as runs are submitted. The Local Intelligence Index (LB-2026-07.2 | 25/22.5/22.5/22.5/7.5) appears only after a complete run publishes. Profile: Agentic / Knowledge / Instruction / Coding / Math.

Every complete project and community run shares this ranking and the same composite. The global view is a cross-family reference; browse model families to choose among related variants. Note: the Agentic axis is near-floor for every current local entrant, so it compresses headline gaps — read the composite alongside the per-axis columns and unweighted diagnostics.

Local Intelligence Index | LB-2026-07.2

Headline = 0.25 Agentic + 0.225 Knowledge + 0.225 Instruction + 0.225 Coding + 0.075 Math. Agentic is AppWorld task-goal completion on the fixed 96-task subset; ranks require the complete headline profile.

5 complete ranked runs

Swipe horizontally for scores and axes →

Complete project and community runs share one score-sorted ranked table.
RankVRAM @8k
1Gemma 4 31B ITLB-2026-07.2
Q4_K_M
Gemma 4
run by local-bench
51.7±3.1
Agentic 10.4 / Knowledge 87.4 / Instruction 79.3 / Coding 35.5 / Math 48.2
10/96 — 10.4%±8.5
diagnostics
Diagnostics · unweighted
Call formatting
74.2%
BFCL single-turn
not measured
BFCL v3 multi-turn base — frozen snapshot
28.0%
BFCL multi-turn long-context
not measured
RULER 32K
not measured
87.4±3.6
79.3±4.8
35.5±7.8
48.2±8.6
20.9 GBllama.cppb9852/fd1a05791RTX 5090 · 32 GB1,918~40 s27.8 h
2Qwen3.6 27BLB-2026-07.2
Q4_K_M
Qwen3.6
run by local-bench
43.2±2.9
Agentic 8.3 / Knowledge 83.4 / Instruction 65.0 / Coding 25.5 / Math 26.6
8/96 — 8.3%±7.4
diagnostics
Diagnostics · unweighted
Call formatting
73.3%
BFCL single-turn
not measured
BFCL v3 multi-turn base — frozen snapshot
22.0%
BFCL multi-turn long-context
not measured
RULER 32K
not measured
83.4±3.9
65.0±5.4
25.5±7.1
26.6±7.2
19.5 GBllama.cppb9852/fd1a05791RTX 5090 · 32 GB3,287~47 s23.7 h
3Qwopus 3.6 27B v2 MTPLB-2026-07.2
Q4_K_M
Qwen3.6
Fine-tune of Qwen3.6 27B
run by local-bench
42.1±3.0
Agentic 9.4 / Knowledge 80.4 / Instruction 57.8 / Coding 29.8 / Math 25.9
9/96 — 9.4%±8.0
diagnostics
Diagnostics · unweighted
Call formatting
73.9%
BFCL single-turn
not measured
BFCL v3 multi-turn base — frozen snapshot
24.0%
BFCL multi-turn long-context
not measured
RULER 32K
not measured
80.4±4.2
57.8±5.8
29.8±7.8
25.9±7.2
18.1 GBllama.cppb9852/fd1a05791RTX 5090 · 32 GB1,623~25 s23.6 h
4Qwen3.6 35B A3BLB-2026-07.2
UD-Q4_K_M
Qwen3.6
run by local-bench
41.0±2.7
Agentic 6.3 / Knowledge 82.0 / Instruction 57.5 / Coding 28.4 / Math 22.3
6/96 — 6.3%±5.4
diagnostics
Diagnostics · unweighted
Call formatting
71.2%
BFCL single-turn
not measured
BFCL v3 multi-turn base — frozen snapshot
20.0%
BFCL multi-turn long-context
not measured
RULER 32K
not measured
82.0±4.2
57.5±5.8
28.4±7.8
22.3±7.2
24.3 GBllama.cppb9852/fd1a05791RTX 5090 · 32 GB3,639~19 s9.2 h
5Gemma 4 12B ITLB-2026-07.2
QAT Q4_K_XL
Gemma 4
run by local-bench
40.3±2.7
Agentic 4.2 / Knowledge 75.6 / Instruction 66.7 / Coding 28.4 / Math 11.5
4/96 — 4.2%±5.8
diagnostics
Diagnostics · unweighted
Call formatting
73.6%
BFCL single-turn
not measured
BFCL v3 multi-turn base — frozen snapshot
24.0%
BFCL multi-turn long-context
not measured
RULER 32K
not measured
75.6±4.8
66.7±5.4
28.4±7.8
11.5±5.8
8.4 GBllama.cppb9852/fd1a05791RTX 5090 · 32 GB8,259~60 s17.2 h

Agentic = AppWorld task-goal completion; 25% weight; near-floor scores compress gaps.

Not yet benchmarked — 159 catalog models on the roadmap