Skip to content

Gemma 4 12B IT

Fine-tune of google/gemma-4-12B

Every measured variant of this model: how much quality each quant keeps, and the VRAM and speed it costs to run.

VRAM footprint vs Local Intelligence Index

LB-2026-07.2 | 25/22.5/22.5/22.5/7.5. Where this model and current-lane family runs land vs the frontier anchors.

Swipe horizontally to inspect the full chart →

10075502508GB12GB16GB24GB32GB48GB64GBQAT Q4_K_XL: 40.3 — benched in 17.2 h · ~6.7 GB to runQAT Q4_K_XLQAT Q4_K_XL — 40.3benched in 17.2 h · ~6.7 GB to runmodel memory footprint (GB)index score4 GB64 GB
This modelDashed vertical lines mark common VRAM tiers.Amber points are synthetic demo preview data.

Variant profiles

Complete rows are ordered by Local Intelligence Index; partial rows show their measured axes but are not ranked. The VRAM/Fits columns (8Kcontext) tell you what your card needs. Ranks are within this family's variants.

Column guide

Variant — quant label plus, where sha-verified, the publisher of the exact weights file benchmarked; rows without a source line predate artifact identity records.

VRAM @8k — model weights, KV cache, and runtime headroom at 8k context.

Fits — the smallest common GPU VRAM tier above that estimate.

Prefill tok/s — prompt-processing speed.

Decode tok/s — generated-token speed after the prompt.

Overall tok/s — completion throughput across the full benchmark.

File size — the benchmarked model artifact on disk.

Runtime — the serving engine and version.

Run — the immutable benchmark receipt. Live rows link to the public submission record until the run is baked into the static site.

Pick the largest quant whose VRAM @8k fits your card.

Swipe horizontally for all variant metrics →

Rank (this family)VariantLocal Intelligence IndexLB-2026-07.2 | 25/22.5/22.5/22.5/7.5AgenticKnowledgeInstructionCodingMathVRAM @8kFitsPrefill tok/sDecode tok/sOverall tok/sFile sizeRuntimeRun
1
QAT Q4_K_XLbestby unsloth
40.3±2.7
Agentic 4.2 / Knowledge 75.6 / Instruction 66.7 / Coding 28.4 / Math 11.5
4.2±5.8
75.6±4.8
66.7±5.4
28.4±7.8
11.5±5.8
6.7 GB footprintn/a5,383.5147.5137.1llama.cppb9852/fd1a05791receipt
QAT Q2_K_XLpartial headlineby unsloth
diagnostic partial — no comparable Index
n/a
23.7±5.1
27.2±5.1
2.1±2.8
0.0±0.0
4.7 GB footprintn/a5,460.4160.5190.1llama.cppb9852/fd1a05791receipt

vs base

Fine-tune comparison

Gemma 4 12B Coder Fable Composer

Fine-tune of Gemma 4 12B IT

composite n/acompare to base
fine-tune not yet benchmarked

Measured axis deltas appear after both rows have board data.

Gemma 4 12B OBLITERATED

Fine-tune of Gemma 4 12B IT

composite n/acompare to base
fine-tune not yet benchmarked

Measured axis deltas appear after both rows have board data.

Gemma 4 12B it AEON Abliterated K4 BF16

Fine-tune of Gemma 4 12B IT

composite n/acompare to base
fine-tune not yet benchmarked

Measured axis deltas appear after both rows have board data.