Skip to content

This receipt lets anyone verify this run against the frozen suite — see Methodology.

← back to Gemma 4 31B IT

suite-v2 | index-v4.2

Gemma 4 31B IT

gemma-4-31b-it__gemma-4-31b-it-q4km-s2v5

Local Intelligence Index
LB-2026-07.2 | 25/22.5/22.5/22.5/7.5
51.7
±3.1 95% CI
Profile: Agentic / Knowledge / Instruction / Coding / Math
Agentic 10.4 / Knowledge 87.4 / Instruction 79.3 / Coding 35.5 / Math 48.2
Weighted headline profile: Agentic 25%, Knowledge 22.5%, Instruction 22.5%, Coding 22.5%, Math 7.5%.
Total run time
27.8 h
Data quality note: this run has 0 error(s) and 7 no-answer item(s).
Data warnings
  • tool_use chance_corrected differs from item-derived mean; stale/inconsistent run JSON
  • coding chance_corrected differs from item-derived mean; stale/inconsistent run JSON

Axis breakdown

Agenticworst axis
n=96 · errors=0 · no answer=0
10.4 ±8.5
Knowledge
n=400 · errors=0 · no answer=1
87.4 ±3.6
Instruction
n=294 · errors=0 · no answer=0
79.3 ±4.8
Coding
n=141 · errors=0 · no answer=0
35.5 ±7.8
Math
n=139 · errors=0 · no answer=6
48.2 ±8.6

Diagnostics · unweighted

Call formatting
n=330 · errors=0 · no answer=0
74.2 ±4.8
BFCL single-turn
— not measured
BFCL v3 multi-turn base — frozen snapshot
n=50 · errors=0 · no answer=0
28.0 ±12.0
BFCL multi-turn long-context
— not measured
RULER 32K
— not measured

Instruction-following decomposition

Strict = Termination × Conditional
Strict accuracy
79.3%
correct AND terminated / all
Termination rate
99.3%
terminated / all
Conditional accuracy
79.8%
correct / terminated

Outputs that hit the answer-token cap are counted incorrect; this prevents non-terminating generations from getting credit for matching required tokens inside a runaway response.

Manifest

model
gemma4
quant
Q4_K_M
runtime
llama.cpp b9852/fd1a05791 · KV k=f16,v=f16 · ctx 32,768
hardware
NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0
os
Windows-11-10.0.26200-SP0
lane
bounded-final-v2
thinking_mode
n/a
caps
max_tokens_math: 0, max_tokens_mcq: 16384, thinking_budget: 8192
sampling
temp 0 | top_p n/a | top_k 1 | min_p n/a | seed 1234 | effort n/a | knowledge: max 16384 | instruction: max 16384 | coding: max 16384 | math: max 16384 | bfcl_multi_turn_base: max 16384 | tc_json_v1: max 16384
tokens
5,527,046 prompt / 4,848,815 completion / 10,375,861 total
tokens-to-answer
1,918 median / 9,318 p95
tok/s
48.4
total run time
27.8 h
est cost
n/a
n_items
1,457
n_errors
0
n_no_answer
7

Serving performance

NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0

prefill
2,425.1 tok/s
decode
63.1 tok/s
TTFT proxy
252.6 ms
prompt processing before first token — non-streaming harness, lower bound
coverage
93.4%
prompt median / p95
252.6 ms / 801.3 ms
predicted median / p95
29,984.3 ms / 152,744.1 ms
benchprefilldecodeprompt mediann
amo818.7 tok/s58 tok/s290.3 ms39
bfcl_multi_turn_base3,089.3 tok/s63.5 tok/s1,659.2 ms50
bigcodebench_hard2,354.9 tok/s59.1 tok/s143.8 ms148
ifbench2,265.3 tok/s66.3 tok/s197.7 ms294
mmlu_pro2,188.9 tok/s62.3 tok/s156.8 ms400
olymmath_hard612.4 tok/s65.1 tok/s255.7 ms100
tc_json_v12,791.1 tok/s60.2 tok/s392.1 ms330

Source: llama.cpp server timings.

Provenance

suite_version: suite-v2

index_version: index-v4.2

source scorecard

version: 6

id: a107306fb9989f8d9b92490c051ebc6a4fb44f0ae62a4a948eda1d0f9a92481b

Provenance drift: this receipt preserves its original scorecard metadata; the site projection is rendered under index-v4.2.

bfcl.jsonl26d990d589db8a8b2a70b23c592ea0aff9df8287d9509a2fadd98d1b72661e17
bfcl_multi_turn_base.jsonl41e5691e87d5c46f12ea19e2bcfe45cb019d83b279c6f1a640bf7ab203c3eec2
bfcl_multi_turn_long_context.jsonlc6f642197ab070573fc353bb7265e7efc813addc27eb59b66358830b68b6a1b2
coding.jsonl33635febb89ab6cb8f06e139bc33932ada89d90e32ce03820ad7f15712e19b8e
instruction.jsonl40dc0b3e14270d61e9deae13f30f70f04d1d65a304340a7b6fe29cf4a5c51257
knowledge.jsonl129b8d9726eab3676ca30d58fac23af4e07407eb537b9bfa10d4d24434b26ba4
math-part2.jsonl8126598901f0e2be27b2a4fed97fded7b2c43aa37ca3ecb580527ad11a15e53b
math.jsonl98e79f1da84680345224f48fc7d1ed8b220e76cfd0525da1c494633d1abd1904
ruler_32k.jsonl0bede1810663a7164e68f3008248d78ba247fb677440f06b3e1c63b8781b0540
tc_json_v1.jsonl571b3c4064b523174900883c786df4fdbb6c2a8924a148620a167415d67afd74