This receipt lets anyone verify this run against the frozen suite — see Methodology.
← back to Qwopus3.6 27B v2 MTPsuite-v2 | index-v4.2
Qwopus3.6 27B v2 MTP
qwopus3-6-27b-v2-mtp__qwopus3-6-27b-v2-mtp-q4km-s2v5
Local Intelligence Index
LB-2026-07.2 | 25/22.5/22.5/22.5/7.5
42.1
±3.0 95% CI
Profile: Agentic / Knowledge / Instruction / Coding / Math
Agentic 9.4 / Knowledge 80.4 / Instruction 57.8 / Coding 29.8 / Math 25.9Weighted headline profile: Agentic 25%, Knowledge 22.5%, Instruction 22.5%, Coding 22.5%, Math 7.5%.
Total run time
23.6 h
Data quality note: this run has 0 error(s) and 49 no-answer item(s).
Data warnings
- coding chance_corrected differs from item-derived mean; stale/inconsistent run JSON
Axis breakdown
Agenticworst axis
n=96 · errors=0 · no answer=0
9.4 ±8.0
Knowledge
n=400 · errors=0 · no answer=12
80.4 ±4.2
Instruction
n=294 · errors=0 · no answer=0
57.8 ±5.8
Coding
n=141 · errors=0 · no answer=0
29.8 ±7.8
Math
n=139 · errors=0 · no answer=37
25.9 ±7.2
Diagnostics · unweighted
Call formatting
n=330 · errors=0 · no answer=1
73.9 ±4.8
BFCL single-turn
— not measured
BFCL v3 multi-turn base — frozen snapshot
n=50 · errors=0 · no answer=0
24.0 ±12.0
BFCL multi-turn long-context
— not measured
RULER 32K
— not measured
Instruction-following decomposition
Strict = Termination × ConditionalStrict accuracy
57.8%
correct AND terminated / all
Termination rate
93.9%
terminated / all
Conditional accuracy
61.6%
correct / terminated
Outputs that hit the answer-token cap are counted incorrect; this prevents non-terminating generations from getting credit for matching required tokens inside a runaway response.
Manifest
- model
- qwen35
- quant
- Q4_K_M
- runtime
- llama.cpp b9852/fd1a05791 · KV k=f16,v=f16 · ctx 32,768
- hardware
- NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0
- os
- Windows-11-10.0.26200-SP0
- lane
- bounded-final-v2
- thinking_mode
- n/a
- caps
- max_tokens_math: 0, max_tokens_mcq: 16384, thinking_budget: 8192
- sampling
- temp 0 | top_p n/a | top_k 1 | min_p n/a | seed 1234 | effort n/a | knowledge: max 16384 | instruction: max 16384 | coding: max 16384 | math: max 16384 | bfcl_multi_turn_base: max 16384 | tc_json_v1: max 16384
- tokens
- 5,690,781 prompt / 5,429,223 completion / 11,120,004 total
- tokens-to-answer
- 1,623 median / 16,384 p95
- tok/s
- 63.8
- total run time
- 23.6 h
- est cost
- n/a
- n_items
- 1,457
- n_errors
- 0
- n_no_answer
- 49
Serving performance
NVIDIA GeForce RTX 5090 (31.8 GB) · Windows-11-10.0.26200-SP0
prefill
2,729.1 tok/s
decode
71 tok/s
TTFT proxy
396.1 ms
prompt processing before first token — non-streaming harness, lower bound
coverage
93.4%
prompt median / p95
396.1 ms / 1,874.1 ms
predicted median / p95
22,347.4 ms / 216,472.4 ms
| bench | prefill | decode | prompt median | n |
|---|---|---|---|---|
| amo | 1,250.6 tok/s | 66.8 tok/s | 197.1 ms | 39 |
| bfcl_multi_turn_base | 3,324.4 tok/s | 72.5 tok/s | 1,976.4 ms | 50 |
| bigcodebench_hard | 2,711.9 tok/s | 66.6 tok/s | 209.7 ms | 148 |
| ifbench | 2,881.6 tok/s | 76.4 tok/s | 381.7 ms | 294 |
| mmlu_pro | 2,582.8 tok/s | 69.5 tok/s | 362.4 ms | 400 |
| olymmath_hard | 945.7 tok/s | 71.6 tok/s | 161.4 ms | 100 |
| tc_json_v1 | 2,657.5 tok/s | 68.7 tok/s | 471.4 ms | 330 |
Source: llama.cpp server timings.