From-scratch language machine

Forty-seven million parameters.Assembled by hand. Zero pretrained weights.

A sparse twist on the small transformer — GQA + RoPE + SwiGLU and Mamba interleaved, crowned by a top-2 Mixture of Experts with a domain-seeded router. Trained under a hard 50M ceiling on a single GPU.

MoE · router activity monitorsweep 1s/div
ACTIVE EXPERTS
2 / 8
ROUTING MODE
TOP-K
GUIDE CURRICULUM
2,000 STEP
SIGNAL
LIVE
AWAITING INPUT…
01 · MODEL

A single track, from token to answer.

Seven interleaved Mamba ∥ Attention blocks feed a top-2 MoE. Every block shares the 50M budget — enforced by two gates.

00 · DATA
EMBED
token lookup · tied head
01 · BLOCK
MB ∥ AT
mamba · attention
MAMBA · SSM ∞GQA · RoPE · SWIGLU
02 · BLOCK
MB ∥ AT
mamba · attention
MAMBA · SSM ∞GQA · RoPE · SWIGLU
03 · BLOCK
MB ∥ AT
mamba · attention
MAMBA · SSM ∞GQA · RoPE · SWIGLU
04 · BLOCK
MB ∥ AT
mamba · attention
MAMBA · SSM ∞GQA · RoPE · SWIGLU
05 · BLOCK
MB ∥ AT
mamba · attention
MAMBA · SSM ∞GQA · RoPE · SWIGLU
06 · BLOCK
MB ∥ AT
mamba · attention
MAMBA · SSM ∞GQA · RoPE · SWIGLU
07 · BLOCK
MB ∥ AT
mamba · attention
MAMBA · SSM ∞GQA · RoPE · SWIGLU
08 · GATE
GATE
pre-MoE norm
09 · MOE
MOE ×8
top-2 · domain-seeded
10 · HEAD
HEAD
lm_head · shared embed
TOTAL
47,640,968
params — 38,793,608 active per token · d=384

Every count is exact, projected and re-checked on the real model before a run is allowed to start.

CONTINUE — THE TRACK KEEPS SPOOLING →
02 · ROUTER

Every token is routed to two of eight experts.

The router never sees the whole net — just a cheap linear projection. A 2,000-step guide curriculum seeds each expert with a domain, then anneals away.

ROUTING PREFERENCE · GATE TIME-SERIES (SIM) LIVE
.0.88
.0.71
.0.72
.0.73
.0.29
.0.30
.0.31
.0.06
.0.72
.0.68
.0.54
.0.58
.0.62
.0.20
.0.24
.0.28
.0.41
.0.63
.0.58
.0.45
.0.50
.0.56
.0.15
.0.21
.0.31
.0.56
.0.81
.0.78
.0.62
.0.65
.0.67
.0.24
.0.49
.0.28
.0.52
.0.76
.0.73
.0.58
.0.61
.0.64
.0.28
.0.54
.0.34
.0.59
.0.85
.0.83
.0.67
.0.68
.0.06
.0.12
.0.30
.0.06
.0.20
.0.38
.0.28
.0.19
.0.21
.0.06
.0.10
.0.27
.0.06
.0.16
.0.33
.0.23
0.001.00
03 · BENCH

Five gates stand between it and competence.

The five mandatory GIBC tasks, run through lm_eval. Baselines shown for context — a 117M GPT-2 and a 70M Pythia.

TASK / METRICMIRALM · 47MGPT-2 · 117MPYTHIA · 70M
HELLASWAG
acc_norm
33.0measured
32.7baseline27–30baseline+0.9%common-sense inference, 4-way MC
ARC-EASY
acc_norm
27.0measured
43.3baseline40–45baseline-37.6%grade-school science, 4-way MC
PIQA
acc_norm
50.0measured
64.2baseline61–63baseline-22.1%physical commonsense, 2-way MC
WINOGRANDE
acc
50.2measured
49.9baseline50–52baseline+0.6%pronoun resolution, 2-way MC
WIKITEXT-103
word_ppl
2837.6measured
~38baseline~44baselineNaN%held-out prose slice; our training corpus is code/math, so this domain shift dominates
5/5 MEASURED Δ = MIRALM vs GPT-2WIKI-103 REPORTED AS WORD PPL · LOWER IS BETTER
04 · LIVE

Ask it in the shape you want back.

Structured output is a protocol: explicit special tokens for JSON, SQL and reasoning. Machine-parseable by construction — no format hacks.

INPUT

> Return a JSON object with id 42, name Alice, age 29, city Kyiv.

activeE05E03
OUTPUT · tok 0/19

Concept render — the real model answers in the web demo.
THE CONTRACT
01

Zero pretrained weights.

Every parameter is randomized and learned on our own corpus. No distillation, no warm-start — the honest way to test a small architecture.

0 · weights reused
02

The budget is enforced, not hoped for.

A two-gate pipeline checks the parameter math at build time — a static projection and a real-model count — and fails the run if the 50M line is crossed.

50M · hard ceiling, CI-enforced
03

Routed by domain, seeded by curriculum.

Eight experts, two active per token. A guide-loss teaches each expert its specialty over the first 2k steps, then anneals away.

8→2 · experts per token
04

One consumer GPU.

fp16 training on a single notebook T4. If your architecture doesn't learn under that constraint, it isn't honest.

1×T4 · fp16 AMP
05

Structured output is a protocol.

JSON / SQL / CoT wrapped in explicit special tokens — machine-parseable answers, extractable reasoning, zero templating hacks.

6 · special tokens
FACTS · ZERO WEIGHTS OUTSIDE THESE LINES
0
total parameters
0
active per token
0
vocabulary size
0
routed experts

No distilled weights. No warm starts. Just a 384-wide net and a patient training loop.

05 · BUILD

Six weeks, one notebook GPU.

The build was serial on purpose: architecture, then data, then training. Each stage shipped with its own test suite before the next began.

  1. W1DONE
    Architecture & the 50M gate
    Block accounting, dual parameter gates, attention core.
  2. W2DONE
    Mamba & Mixture of Experts
    Selective SSM, top-2 router, guide curriculum.
  3. W3DONE
    Data & the 24k tokenizer
    ByteLevel BPE, packing, memmap shards, domains.
  4. W4IN TRAINING
    Training loop
    fp16 AMP, cosine LR, checkpointing, W&B.
  5. W5OPEN
    Structured SFT
    JSON / SQL / CoT protocol on top of the pretrained stack.
  6. W6OPEN
    Evaluation & demo
    Five mandatory benchmarks, web demo, submission.