╱/ Slayer BDH Benchmark Ladder

progres.fabryka.ai · Dragon Hatchling scaling series · byte-level (vocab 256) · live benchmark results

Param scope: Every model is scored only on rungs where it has resolution above chance (L0 = continuous signal; L1 at ≥20–50M; L2 at ≥100–300M; L3 needs 1B+, out of scope). Scores below min_resolution delta-from-chance are flagged NO-SIGNAL.

L0 Held-out Language Quality (Bits-Per-Byte)  · always run

ModelStepsBPB ∕ ENBPB ∕ PLEN⇄PL gapStatus

L0 Associative Recall (fast-weights / in-context thesis)

ModelVAccChanceDeltaSignal

BDH-specific probe: can the model recall value from key seen earlier in context? at  chance → fast-weights not yielding in-context memory; above chance → thesis holds.

L1 Language / Commonsense MCQ

BenchLang50M150M350MChance

L2 Reasoning-rungs (needs ≥100–300M)

BenchLang50M150M350MChance

L3 Knowledge / Math (needs 1B+ — OUT OF SCOPE)

BenchChanceNote