Skip to content

This receipt lets anyone verify this run against the frozen suite — see Methodology.

← back to Gemma 4 12B IT

suite-v2 | index-v4.2

Gemma 4 12B IT

gemma-4-12b-it__gemma-4-12b-it-qat-ud-q4kxl-s2v5

Local Intelligence Index
LB-2026-07.2 | 25/22.5/22.5/22.5/7.5
40.3
±2.7 95% CI
Profile: Agentic / Knowledge / Instruction / Coding / Math
Agentic 4.2 / Knowledge 75.6 / Instruction 66.7 / Coding 28.4 / Math 11.5
Weighted headline profile: Agentic 25%, Knowledge 22.5%, Instruction 22.5%, Coding 22.5%, Math 7.5%.
Total run time
17.2 h
Data quality note: this run has 0 error(s) and 103 no-answer item(s).
Data warnings
  • coding chance_corrected differs from item-derived mean; stale/inconsistent run JSON

Axis breakdown

Agenticworst axis
n=96 · errors=0 · no answer=0
4.2 ±5.8
Knowledge
n=400 · errors=0 · no answer=16
75.6 ±4.8
Instruction
n=294 · errors=0 · no answer=0
66.7 ±5.4
Coding
n=141 · errors=0 · no answer=0
28.4 ±7.8
Math
n=139 · errors=0 · no answer=87
11.5 ±5.8

Diagnostics · unweighted

Call formatting
n=330 · errors=0 · no answer=0
73.6 ±4.8
BFCL single-turn
— not measured
BFCL v3 multi-turn base — frozen snapshot
n=50 · errors=0 · no answer=6
24.0 ±12.0
BFCL multi-turn long-context
— not measured
RULER 32K
— not measured

Instruction-following decomposition

Strict = Termination × Conditional
Strict accuracy
66.7%
correct AND terminated / all
Termination rate
95.9%
terminated / all
Conditional accuracy
69.5%
correct / terminated

Outputs that hit the answer-token cap are counted incorrect; this prevents non-terminating generations from getting credit for matching required tokens inside a runaway response.

Manifest

model
gemma4
quant
QAT Q4_K_XL
runtime
llama.cpp b9852/fd1a05791 · KV k=f16,v=f16 · ctx 32,768
hardware
NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0
os
Windows-11-10.0.26200-SP0
lane
bounded-final-v2
thinking_mode
n/a
caps
max_tokens_math: 0, max_tokens_mcq: 16384, thinking_budget: 8192
sampling
temp 0 | top_p n/a | top_k 1 | min_p n/a | seed 1234 | effort n/a | knowledge: max 16384 | instruction: max 16384 | coding: max 16384 | math: max 16384 | bfcl_multi_turn_base: max 16384 | tc_json_v1: max 16384
tokens
8,355,858 prompt / 8,468,917 completion / 16,824,775 total
tokens-to-answer
8,259 median / 16,384 p95
tok/s
137.1
total run time
17.2 h
est cost
n/a
n_items
1,457
n_errors
0
n_no_answer
103

Serving performance

NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0

prefill
5,383.5 tok/s
decode
147.5 tok/s
TTFT proxy
114.7 ms
prompt processing before first token — non-streaming harness, lower bound
coverage
93.4%
prompt median / p95
114.7 ms / 978.5 ms
predicted median / p95
54,977.8 ms / 109,719.3 ms
benchprefilldecodeprompt mediann
amo4,022.3 tok/s149.6 tok/s120.6 ms39
bfcl_multi_turn_base6,499.7 tok/s128.8 tok/s754.1 ms50
bigcodebench_hard6,073.7 tok/s151.1 tok/s103.7 ms148
ifbench5,203.1 tok/s145.5 tok/s106.2 ms294
mmlu_pro4,297.9 tok/s150.6 tok/s113.6 ms400
olymmath_hard2,349.8 tok/s145.6 tok/s117.1 ms100
tc_json_v16,926.3 tok/s152.5 tok/s125.5 ms330

Source: llama.cpp server timings.

Provenance

suite_version: suite-v2

index_version: index-v4.2

source scorecard

version: 6

id: a107306fb9989f8d9b92490c051ebc6a4fb44f0ae62a4a948eda1d0f9a92481b

Provenance drift: this receipt preserves its original scorecard metadata; the site projection is rendered under index-v4.2.

bfcl.jsonl26d990d589db8a8b2a70b23c592ea0aff9df8287d9509a2fadd98d1b72661e17
bfcl_multi_turn_base.jsonl41e5691e87d5c46f12ea19e2bcfe45cb019d83b279c6f1a640bf7ab203c3eec2
bfcl_multi_turn_long_context.jsonlc6f642197ab070573fc353bb7265e7efc813addc27eb59b66358830b68b6a1b2
coding.jsonl33635febb89ab6cb8f06e139bc33932ada89d90e32ce03820ad7f15712e19b8e
instruction.jsonl40dc0b3e14270d61e9deae13f30f70f04d1d65a304340a7b6fe29cf4a5c51257
knowledge.jsonl129b8d9726eab3676ca30d58fac23af4e07407eb537b9bfa10d4d24434b26ba4
math-part2.jsonl8126598901f0e2be27b2a4fed97fded7b2c43aa37ca3ecb580527ad11a15e53b
math.jsonl98e79f1da84680345224f48fc7d1ed8b220e76cfd0525da1c494633d1abd1904
ruler_32k.jsonl0bede1810663a7164e68f3008248d78ba247fb677440f06b3e1c63b8781b0540
tc_json_v1.jsonl571b3c4064b523174900883c786df4fdbb6c2a8924a148620a167415d67afd74