Skip to content

This receipt lets anyone verify this run against the frozen suite — see Methodology.

← back to Qwen3.6-27B

suite-v2 | index-v4.2

Qwen3.6-27B

qwen3-6-27b__qwen3-6-27b-q4km-s2v5

Local Intelligence Index
LB-2026-07.2 | 25/22.5/22.5/22.5/7.5
43.2
±2.9 95% CI
Profile: Agentic / Knowledge / Instruction / Coding / Math
Agentic 8.3 / Knowledge 83.4 / Instruction 65.0 / Coding 25.5 / Math 26.6
Weighted headline profile: Agentic 25%, Knowledge 22.5%, Instruction 22.5%, Coding 22.5%, Math 7.5%.
Total run time
23.7 h
Data quality note: this run has 0 error(s) and 10 no-answer item(s).
Data warnings
  • tool_use chance_corrected differs from item-derived mean; stale/inconsistent run JSON
  • coding chance_corrected differs from item-derived mean; stale/inconsistent run JSON

Axis breakdown

Agenticworst axis
n=96 · errors=0 · no answer=0
8.3 ±7.4
Knowledge
n=400 · errors=0 · no answer=1
83.4 ±3.9
Instruction
n=294 · errors=0 · no answer=0
65.0 ±5.4
Coding
n=141 · errors=0 · no answer=0
25.5 ±7.1
Math
n=139 · errors=0 · no answer=9
26.6 ±7.2

Diagnostics · unweighted

Call formatting
n=330 · errors=0 · no answer=0
73.3 ±4.8
BFCL single-turn
— not measured
BFCL v3 multi-turn base — frozen snapshot
n=50 · errors=0 · no answer=0
22.0 ±12.0
BFCL multi-turn long-context
— not measured
RULER 32K
— not measured

Instruction-following decomposition

Strict = Termination × Conditional
Strict accuracy
65.0%
correct AND terminated / all
Termination rate
95.6%
terminated / all
Conditional accuracy
68.0%
correct / terminated

Outputs that hit the answer-token cap are counted incorrect; this prevents non-terminating generations from getting credit for matching required tokens inside a runaway response.

Manifest

model
qwen35
quant
Q4_K_M
runtime
llama.cpp b9852/fd1a05791 · KV k=f16,v=f16 · ctx 32,768
hardware
NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0
os
Windows-11-10.0.26200-SP0
lane
bounded-final-v2
thinking_mode
n/a
caps
max_tokens_math: 0, max_tokens_mcq: 16384, thinking_budget: 8192
sampling
temp 0 | top_p n/a | top_k 1 | min_p n/a | seed 1234 | effort n/a | knowledge: max 16384 | instruction: max 16384 | coding: max 16384 | math: max 16384 | bfcl_multi_turn_base: max 16384 | tc_json_v1: max 16384
tokens
6,525,242 prompt / 5,921,541 completion / 12,446,783 total
tokens-to-answer
3,287 median / 9,380 p95
tok/s
69.4
total run time
23.7 h
est cost
n/a
n_items
1,457
n_errors
0
n_no_answer
10

Serving performance

NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0

prefill
3,112.1 tok/s
decode
74.4 tok/s
TTFT proxy
640.9 ms
prompt processing before first token — non-streaming harness, lower bound
coverage
93.4%
prompt median / p95
640.9 ms / 2,249.7 ms
predicted median / p95
43,212.2 ms / 130,836.7 ms
benchprefilldecodeprompt mediann
amo1,412 tok/s76.4 tok/s171.3 ms39
bfcl_multi_turn_base3,400.4 tok/s74.3 tok/s2,500.5 ms50
bigcodebench_hard3,345.6 tok/s77.1 tok/s1,306.4 ms148
ifbench3,163.1 tok/s75.5 tok/s935.5 ms294
mmlu_pro2,948.9 tok/s72.8 tok/s572.1 ms400
olymmath_hard874.5 tok/s72.5 tok/s169.2 ms100
tc_json_v13,115.1 tok/s77.8 tok/s566.5 ms330

Source: llama.cpp server timings.

Provenance

suite_version: suite-v2

index_version: index-v4.2

source scorecard

version: 6

id: d7ba660a32a9abf9a3e6ef469f3e33e2c6b655de9331e3af7a4fed493fd0491c

Provenance drift: this receipt preserves its original scorecard metadata; the site projection is rendered under index-v4.2.

bfcl.jsonl26d990d589db8a8b2a70b23c592ea0aff9df8287d9509a2fadd98d1b72661e17
bfcl_multi_turn_base.jsonl41e5691e87d5c46f12ea19e2bcfe45cb019d83b279c6f1a640bf7ab203c3eec2
bfcl_multi_turn_long_context.jsonlc6f642197ab070573fc353bb7265e7efc813addc27eb59b66358830b68b6a1b2
coding.jsonl33635febb89ab6cb8f06e139bc33932ada89d90e32ce03820ad7f15712e19b8e
instruction.jsonl40dc0b3e14270d61e9deae13f30f70f04d1d65a304340a7b6fe29cf4a5c51257
knowledge.jsonl129b8d9726eab3676ca30d58fac23af4e07407eb537b9bfa10d4d24434b26ba4
math-part2.jsonl8126598901f0e2be27b2a4fed97fded7b2c43aa37ca3ecb580527ad11a15e53b
math.jsonl98e79f1da84680345224f48fc7d1ed8b220e76cfd0525da1c494633d1abd1904
ruler_32k.jsonl0bede1810663a7164e68f3008248d78ba247fb677440f06b3e1c63b8781b0540
tc_json_v1.jsonl571b3c4064b523174900883c786df4fdbb6c2a8924a148620a167415d67afd74