Forty-seven million parameters.Assembled by hand. Zero pretrained weights.
A sparse twist on the small transformer — GQA + RoPE + SwiGLU and Mamba interleaved, crowned by a top-2 Mixture of Experts with a domain-seeded router. Trained under a hard 50M ceiling on a single GPU.
A single track, from token to answer.
Seven interleaved Mamba ∥ Attention blocks feed a top-2 MoE. Every block shares the 50M budget — enforced by two gates.
Every count is exact, projected and re-checked on the real model before a run is allowed to start.
Every token is routed to two of eight experts.
The router never sees the whole net — just a cheap linear projection. A 2,000-step guide curriculum seeds each expert with a domain, then anneals away.
Five gates stand between it and competence.
The five mandatory GIBC tasks, run through lm_eval. Baselines shown for context — a 117M GPT-2 and a 70M Pythia.
Ask it in the shape you want back.
Structured output is a protocol: explicit special tokens for JSON, SQL and reasoning. Machine-parseable by construction — no format hacks.
> Return a JSON object with id 42, name Alice, age 29, city Kyiv.
Zero pretrained weights.
Every parameter is randomized and learned on our own corpus. No distillation, no warm-start — the honest way to test a small architecture.
0 · weights reusedThe budget is enforced, not hoped for.
A two-gate pipeline checks the parameter math at build time — a static projection and a real-model count — and fails the run if the 50M line is crossed.
50M · hard ceiling, CI-enforcedRouted by domain, seeded by curriculum.
Eight experts, two active per token. A guide-loss teaches each expert its specialty over the first 2k steps, then anneals away.
8→2 · experts per tokenOne consumer GPU.
fp16 training on a single notebook T4. If your architecture doesn't learn under that constraint, it isn't honest.
1×T4 · fp16 AMPStructured output is a protocol.
JSON / SQL / CoT wrapped in explicit special tokens — machine-parseable answers, extractable reasoning, zero templating hacks.
6 · special tokensNo distilled weights. No warm starts. Just a 384-wide net and a patient training loop.
Six weeks, one notebook GPU.
The build was serial on purpose: architecture, then data, then training. Each stage shipped with its own test suite before the next began.
- W1DONEArchitecture & the 50M gateBlock accounting, dual parameter gates, attention core.
- W2DONEMamba & Mixture of ExpertsSelective SSM, top-2 router, guide curriculum.
- W3DONEData & the 24k tokenizerByteLevel BPE, packing, memmap shards, domains.
- W4IN TRAININGTraining loopfp16 AMP, cosine LR, checkpointing, W&B.
- W5OPENStructured SFTJSON / SQL / CoT protocol on top of the pretrained stack.
- W6OPENEvaluation & demoFive mandatory benchmarks, web demo, submission.