LB-2026-07.2 | scorecard-v6 methodology
How local-bench scores runs
The sortable number is the Local Intelligence Index (LB-2026-07.2 | 25/22.5/22.5/22.5/7.5): 25% Agentic, 22.5% Knowledge, 22.5% Instruction-Following, 22.5% Coding, and 7.5% Math. Profile: Agentic / Knowledge / Instruction / Coding / Math stays visible beside every score. Season-1 (index-v3.0) results remain published as history and diagnostics; the two scales are never directly compared — see the season bridge below.
Glossary
- VRAM @8k
- Estimated GPU memory for model weights, an 8k-token KV cache, and runtime headroom.
- Fits
- The smallest common GPU VRAM tier that is at least as large as the VRAM @8k estimate.
- Prefill tok/s
- How quickly the runtime reads and processes the prompt before generating an answer.
- Decode tok/s
- How quickly the runtime generates new answer tokens after processing the prompt.
- Tokens/answer
- The median number of generated tokens per answer, a practical measure of verbosity.
- Wall/bench time
- The elapsed time to finish the full benchmark run, including every measured axis.
- Lane
- A fixed serving-engine and benchmark-protocol configuration used to keep rows comparable.
- Headline profile
- The complete required set of weighted axes a run must publish before it can receive a rank.
Protocol v4.2 · LB-2026-07.2
One scoring protocol for every ranked row
Community and project rows must be scored identically. The previous Agentic axis blended AppWorld with a benchmark that community runs could not execute, so the two populations were ranked on unequal compositions. Protocol v4.2 corrects that comparability defect by making Agentic AppWorld-only everywhere. The structural key remains tool_use, and its 25% headline weight does not change.
suite-v1-full-exec-6axis-v1 is the current ranked suite. The suite measures six axes; five are weighted in the Index, tool-calling is reported as an unweighted diagnostic.
Agentic measures AppWorld task-goal completion under the published runner: a fixed 96-task subset from test_normal, selected with seeded stratified recipe v1 and seed 20260624. The subset is identical for every row. The published protocol exposes counts and hashes, never upstream task contents.
- Subset sha256
- 7aabcf2af32300cf8769ce63cdc09353e7eab3a8681386d46fb0747950c85095
- Ordered task-ID sha256
- 16394237615c06aa419135cca2543ab66957a6af092e7da817565d3518ee648c
- Selection recipe sha256
- b785b74c00d7e24b48826a04003b198f123ecd4fc948d0670b3215984111e3e0
- Protocol manifest sha256
- 45416b98cd17b691e9858a02cc74b510d2b8a16aa5d2585d4672212f8044a416
Correction impact
Raw inference outputs are unchanged; no model was re-run. Existing item verdicts were re-scored under the equal v4.2 protocol. The rank order is unchanged.
Swipe horizontally for v4.1 and v4.2 scores →
| Model | v4.1 | v4.2 |
|---|---|---|
| Gemma 4 31B IT | 53.12 | 51.69 |
| Qwen3.6 27B | 44.35 | 43.22 |
| Qwopus 3.6 27B v2 MTP | 43.27 | 42.08 |
| Qwen3.6 35B A3B | 42.12 | 41.01 |
| Gemma 4 12B IT | 42.03 | 40.29 |
Archived v4.1 snapshots: board JSON and Agentic JSON.
Unweighted diagnostics
Call formatting, BFCL single-turn, BFCL v3 multi-turn base — frozen snapshot, BFCL multi-turn long-context, RULER 32K are displayed when measured and shown as not measured when absent. They are never weighted, never zero-filled, and never used for ranking. BFCL v3 multi-turn base is retained as a frozen, version-pinned diagnostic.
Season 1 → 2 bridge
Index-v3.0 and index-v4.x composites are different editorial scales and are never directly compared. Index-v4.0, index-v4.1, and LB-2026-07.2 are also distinct protocol snapshots. Lineage deltas and ordinary compare views require matching index versions. The versioned season bridge is the only sanctioned pairing: for an artifact measured completely in both seasons, it presents the frozen v3 composite beside the full v4 composite without treating their numerical gap as a performance delta.
An anchor without complete season-2 coverage keeps its season-1 label and composite. A partial v4 composite is never displayed or ranked.
What index-v3.0 measures
The Index is a weighted arithmetic mean of measured, judge-free axes. Agentic is AppWorld-C task success rate. Knowledge is MMLU-Pro. Instruction is IFBench. Tool calling is tc_json_v1 structural tool selection and argument construction. Coding is BigCodeBench-Hard Instruct execution pass rate. Math combines OlymMATH-Hard and AMO. Season-1 headline weights were 40% Agentic, 15% Knowledge, 15% Instruction-Following, 10% Tool calling, 15% Coding, and 5% Math.
The static ranked Index removes Agentic and reweights the remaining five axes: 25% Knowledge, 25% Instruction-Following, 20% Tool calling, 20% Coding, and 10% Math. Static-Core measures only Knowledge, Instruction, Tool calling, and Math; it has no sandbox, no Agentic, and no executable Coding, so it is an unranked diagnostic release and is not comparable to ranked static.
Why the Index is not a parameter-count score
The Local Intelligence Index is a fixed, published weighted average of six local-use capabilities measured under the same generated-token budget — not a proxy for model size. Larger models often gain on knowledge, math, and agentic tasks, but smaller instruction-tuned models can legitimately close the headline gap by following instructions better, producing stronger executable code, or tying on tool-call structure. A 12B model landing within a few points of a 27B is expected behavior on this index, and the per-axis columns show exactly where each model earns or loses its score.
Agentic headroom is deliberate. Low agentic scores on today's board are expected: this axis is not rescaled upward to make current models look more spread out, so future local models that are genuinely capable agents have room to separate rather than saturating the benchmark.
The agentic budget cap is a benchmark restriction, not a usage claim. Agentic tasks run under the same bounded token budget and per-task time limit as every other axis. That is a constraint we impose for comparability and feasible run times, not a claim that local agents need to be token-efficient — at home, letting a model think longer is essentially free. A model cut off by the budget scores the same as one that had no idea, so thinking-heavy models may perform better in unbounded local use than their agentic score here suggests.
Rows whose agentic phase shows an elevated infrastructure timeout rate (over 5 percent) do not rank until re-run on hardware that clears it.
Tool calling scope note. The tool-calling axis measures structural JSON tool selection and argument construction (tc_json_v1), not full real-world tool-use competence. The first ranked rows show low spread on this axis; a harder versioned itemset is planned, and v1 history will be preserved when it lands.
Lane and ranking rules
Ranked rows run the bounded-final lane. Each item gets the same total generated-token budget for that item; optional thinking is force-closed inside that budget, and only the final answer is scored. Code items carry a larger final-answer reserve of 4096 tokens, while ordinary bounded-final items default to 1024.
Execution profiles replace family gating. Eligibility is audits, conformance, and an allowlisted profile digest, never the model family. Legacy v1-lane rows keep their lane and index labels until they are re-run; the default board shows only the current index identity.
The “best at its size” tag and the dotted line on the front-page chart mark the size-vs-score Pareto frontier: a model is on it when no measured model is both higher-scoring and smaller (by the benchmarked artifact’s on-disk size plus estimated KV cache). It is a value-per-VRAM marker computed from point estimates, not a capability tier. The chart shows one point per weights family — the best measured variant across a base model, its fine-tunes, and their quants — while each model page compares variants within the family.
Coding execution and trust
Every publishable bundle includes Coding results produced on the submitter's machine in the pinned, network-isolated local-bench sandbox. Coding and Agentic evidence travel with the bundle; the publication service validates the complete projection and does not re-run model-generated code.
The Coding axis reports pass rate over the 141 sandbox-scoreable BigCodeBench-Hard items; seven network/data-dependent upstream items are excluded as unscoreable under mandatory network isolation. lcb, the old LiveCodeBench output-prediction proxy, is legacy diagnostic data and is never pooled into index-v3.0.
What publication means
Community-reported results publish immediately after the complete headline profile, suite pins, schema, size limits, and duplicate-retry checks pass. The site preserves the submitted identity, protocol, scores, and evidence bundle, computes the common composite, and suppresses rows when problems are demonstrated. Results are not independently reproduced by default.
- One provenance badge. Only rows whose server-owned origin is project_anchor receive the project run badge. Community reports are the unmarked default.
- Names are details, not identity proof. An optional free-text handle appears only as “submitted as … — unverified”. It never replaces the model or artifact identity.
- One ranking rule. Every published row with the complete headline profile enters the same score order. Incomplete legacy records remain available on family pages as history, never as partial board rows.
- Moderation happens after publication. Evidence-backed problems can trigger a consistency check or independent re-run, and the maintainer suppresses a row when the evidence warrants it. There is no pending truth judgment or promotion tier.
Evidence and reproduction
Each row keeps its structured model artifact identity, immutable bundle hash, protocol and suite identity, axis scores, sample counts, confidence intervals, and downloadable evidence. Anyone can use localbench verify for a consistency check or independently re-run the public suite; those are different claims, and the site does not call a consistency check a reproduction.
Threat posture, honestly
An OpenAI-compatible endpoint can be a cheat proxy: it can look up public answers and fabricate plausible transcripts, and no server-side check proves which model — or whose hardware — actually produced a bundle. Public evidence and family-page outliers make scrutiny practical, but they do not prove model identity or honest execution. The policy is publish complete reports, audit consequential leaders or evidence-backed disputes, and suppress demonstrated problems.
Determinism and reproducibility
Scoring is deterministic from frozen artifacts and submitted outputs. Model reruns are reported with fixed settings and bootstrap confidence intervals, not claimed bit-identical across hardware or software stacks.
Every displayed score carries a bootstrap confidence interval. Repeatability, paired quant-delta, and generalization are kept separate. Per-axis confidence intervals are part of the score display; coding deltas under about 8-10 raw points are not ranking claims unless rank containment and intervals support that read. Incomplete historical runs remain diagnostics. The current board admits only rows with the complete headline profile.
Serving engine lanes
Rows identify the serving engine as well as the model format. Community-submitted llama.cpp runs use a pinned GGUF artifact. Safetensors/NVFP4 rows use the project-operated vLLM path in WSL2; that runtime is not yet a supported community provisioning path and remains so until the appliance ships.
A vLLM receipt pins the Hugging Face repository and full 40-character revision, the snapshot Merkle identity and per-file hashes, the server-reported engine version and dependency identity, a two-start determinism canary with engine-log evidence, and the declared model, KV-cache, and Mamba SSM-state dtypes. Its reproduction form is localbench bench --runtime vllm --model-ref hf://<repo>@<revision> --hf-model-id <repo> (run localbench cache-tokenizer <repo> first; the execution profile enforces a context floor of 26624).
Every vLLM row runs under a named determinism policy, shown on the board next to the engine. batch-invariant rows run with VLLM_BATCH_INVARIANT=1 and claim best-effort same-stack reproducibility. GDN/linear-attention hybrids (for example Qwen3.6), which vLLM refuses to initialise in batch-invariant mode, run instead under GDN graphs v2: a narrower, honestly scoped claim — empirically reproducible across clean process starts under structural single-slot execution with a pinned cudagraph configuration and a pinned per-run Triton autotune manifest on the recorded stack. It is not batch-invariant and not cross-stack bitwise deterministic. Kernel autotuning selects winners by wall-clock benchmarking, so the run pins them: the first canary start derives the winners and publishes the manifest hash (plus the full start-1 log hash), and the scoring server replays exactly those winners. A fresh reproduction on other hardware — or even a fresh cold start — may legitimately tune different winners and produce different bytes; the published evidence discloses the pin so that difference is inspectable rather than hidden.
The canary itself is token-level and long-context aware: probes target rendered lengths of 128, 64 and 65 (straddling the GDN chunk boundary), 8k, 16k, 26624, and near the configured context — each within ±2 tokens of target — with within-lifetime repeats, an A/B/A state-isolation re-probe, at most one bounded relaunch of the scoring start, and post-score sentinel re-probes against the still-live scoring server. Provenance additionally requires zero Triton JIT compilations during the scored phase.
Contamination and public items
Several inputs are public benchmarks, so local-bench does not pretend contamination is impossible. The item subsets are frozen and sha256-pinned, contamination status is shown per run, and small leaderboard gaps should be treated as uncertainty unless they clear the reported confidence intervals.
Quant degradation is drift, not just accuracy
Model pages report quantization as a trade-off across accuracy, VRAM, speed, tokens-to-answer, and output drift against a full-precision or explicitly labeled proxy reference. Drift is labeled as drift, never as a hidden task score.
Frozen as of 2026-07-20
The board is a point-in-time snapshot. Run receipts carry their suite, lane, scorecard, and item-set hashes; legacy receipts keep their original labels until they are re-run under the current release.
- As-of date
- 2026-07-20
- Index identity
- LB-2026-07.2 / scorecard-v6
- Board sha256
- c2687d46ac442415641d0478257f3c6d16d583f7113ed8481b2a5a3e1a459a3c
Benchmark item licenses and scorer attribution are listed in the next section.
Benchmark sources & licenses
local-bench publishes public suite artifacts, scores, and metadata. No model weights, GGUF files, API keys, or private runtime credentials are hosted here. The frozen item sets behind the headline Index are derived from the following sources, each used under its own license:
AppWorld-C
Agentic task-goal completion axis (96-task fixed subset).
AppWorld authors
Apache-2.0
MMLU-Pro
Knowledge axis item set (400 items).
TIGER-Lab
MIT
IFBench
Instruction-Following axis item set (294 items).
Allen Institute for AI (Ai2)
ODC-BY-1.0 (dataset)
IFEval checker
Instruction-following verifier logic adapted into the local-bench scorer.
Google Research
Apache-2.0
BigCodeBench-Hard Instruct
Coding axis generation tasks (141 sandbox-scoreable items) scored by hardened execution.
BigCodeBench authors
Apache-2.0
OlymMATH-Hard
Math axis hard olympiad-style item set (100 items).
OlymMATH authors
MIT
AMO
Math axis newly-authored olympiad-style item set (39 items; Math 139 total).
AMO-Bench authors
MIT
Candidate and opt-in modules draw on additional sources, credited the same way:
TC-JSON v1
Unweighted call-formatting diagnostic and structural JSON scorer.
local-bench + Gorilla LLM / UC Berkeley
Apache-2.0
LiveCodeBench / RULER / BFCL expansions
Legacy or candidate diagnostic modules credited in their suite manifests.
LiveCodeBench authors / NVIDIA / Gorilla LLM and UC Berkeley
various open licenses
Full license texts and the complete redistribution notice ship in the repository’s NOTICE file and LICENSES/ directory. Model names are the property of their respective owners; listing a model is benchmark evaluation, not an endorsement by, or of, its maker.
Editorial versioning
Domain weights are explicit editorial choices tied to named releases: LB-2026-07.2 (scorecard-v6) is the current five-axis Index with AppWorld-only Agentic. Index-v4.1 retains the same headline weights but its unequal Agentic composition is archived; index-v4.0 is the initial season-2 scale, and index-v3.0 is the season-1 six-axis Index. Static-suite-v2 and static-core remain non-rankable diagnostics; neither can produce an active board row. Weights and membership live in the versioned protocol manifest and scorer registry, so history cannot be silently re-scored under the same label.