Skip to content

This receipt lets anyone verify this run against the frozen suite — see Methodology.

← back to Gemma 4 12B IT

suite-v1 | index-v3.0

Gemma 4 12B IT

gemma-4-12b-it__gemma-4-12b-it-qat-ud-q2kxl-bounded-final-v2

Diagnostic score profile
index-v3.0 | 40/15/15/10/15/5
17.3
±4.0 95% CI
Profile: Agentic / Knowledge / Instruction / Tool calling / Coding / Math
Agentic n/a / Knowledge 23.7 / Instruction 27.2 / Tool 24.5 / Coding 2.1 / Math 0.0
Weighted headline profile: Agentic 40%, Knowledge 15%, Instruction 15%, Tool 10%, Coding 15%, Math 5%.
Total run time
19.3 h
Data quality note: this run has 3 error(s) and 625 no-answer item(s).
Unranked diagnostic: this receipt is missing one or more headline axes for the current index. It is useful for comparing measured axes, but it does not receive a ranked Local Intelligence Index position.
Data warnings
  • agentic appworld_c quarantined: missing diagnostics or harness-dominated client path
  • long_context not measured in this run; axis omitted
  • agentic not measured in this run; axis omitted

Axis breakdown

Agentic
— not measured
Knowledge
n=400 · errors=3 · no answer=208
23.7 ±5.1
Instruction
n=294 · errors=0 · no answer=0
27.2 ±5.1
Tool calling
n=330 · errors=0 · no answer=204
24.5 ±4.5
Coding
n=141 · errors=0 · no answer=75
2.1 ±2.8
Mathworst axis
n=139 · errors=0 · no answer=138
0.0 ±0.0

Instruction-following decomposition

Strict = Termination × Conditional
Strict accuracy
27.2%
correct AND terminated / all
Termination rate
72.8%
terminated / all
Conditional accuracy
37.4%
correct / terminated

Outputs that hit the answer-token cap are counted incorrect; this prevents non-terminating generations from getting credit for matching required tokens inside a runaway response.

Manifest

model
gemma4
quant
QAT Q2_K_XL
runtime
llama.cpp b9852/fd1a05791 · KV k=f16,v=f16 · ctx 32,768
hardware
NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0
os
Windows-11-10.0.26200-SP0
lane
bounded-final-v2
thinking_mode
n/a
caps
max_tokens_math: 0, max_tokens_mcq: 16384, thinking_budget: 8192
sampling
temp 0 | top_p n/a | top_k 1 | min_p n/a | seed 1234 | effort n/a | knowledge: max 16384 | instruction: max 16384 | tool_calling: max 16384 | coding: max 16384 | math: max 16384
tokens
9,335,171 prompt / 13,209,149 completion / 22,544,320 total
tokens-to-answer
8,618 median / 16,384 p95
tok/s
190.1
total run time
19.3 h
est cost
n/a
n_items
1,407
n_errors
3
n_no_answer
625

Serving performance

NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0

prefill
5,460.4 tok/s
decode
160.5 tok/s
TTFT proxy
174.6 ms
prompt processing before first token — non-streaming harness, lower bound
coverage
93.0%
prompt median / p95
174.6 ms / 1,516.3 ms
predicted median / p95
54,006 ms / 107,746 ms
benchprefilldecodeprompt mediann
amo5,213.7 tok/s141 tok/s312.2 ms39
bigcodebench_hard5,145 tok/s155.8 tok/s146.5 ms148
ifbench5,714.6 tok/s168.3 tok/s133.1 ms294
mmlu_pro5,700.1 tok/s166.7 tok/s154.8 ms397
olymmath_hard4,946.8 tok/s147.7 tok/s140.5 ms100
tc_json_v15,050.4 tok/s156.4 tok/s193.2 ms330

Source: llama.cpp server timings.

Provenance

suite_version: suite-v1

index_version: index-v3.0

source scorecard

version: 3

id: 39edac7728bdff434b56074d7b5b8538f4e4619dc39cb4a175058968ea26e53c

Provenance drift: this receipt preserves its original scorecard metadata; the site projection is rendered under index-v3.0.

coding.jsonl33635febb89ab6cb8f06e139bc33932ada89d90e32ce03820ad7f15712e19b8e
instruction.jsonl40dc0b3e14270d61e9deae13f30f70f04d1d65a304340a7b6fe29cf4a5c51257
knowledge.jsonl129b8d9726eab3676ca30d58fac23af4e07407eb537b9bfa10d4d24434b26ba4
math-part5.jsonl8126598901f0e2be27b2a4fed97fded7b2c43aa37ca3ecb580527ad11a15e53b
math.jsonl98e79f1da84680345224f48fc7d1ed8b220e76cfd0525da1c494633d1abd1904
tool_calling.jsonl571b3c4064b523174900883c786df4fdbb6c2a8924a148620a167415d67afd74