Sane Labs Synth-2

Notes from the run.

Synth-2 is in pretraining, which means most weeks produce a measurement rather than a release. These are the measurements that changed what happened next — including the ones that killed a favourite explanation.

—

Working notes

Note 0410 Sep 2026

The site spent a day in quirks mode

Not a training note, but it cost a day, so it is written down. This page and the two beside it were drafted inside a preview that supplies its own document wrapper. Published to a plain host, the wrapper was gone and so was the doctype — and a document with no doctype is in quirks mode, where document.scrollingElement is <body> rather than <html>. The smooth-scroll library went on writing to documentElement.scrollTop, which in that mode scrolls nothing at all. The page rendered perfectly and would not move.

The second half of the same afternoon: body{overflow-x:hidden} promotes <body> to the scroll container by itself, with the same consequence. overflow-x:clip trims the overflow without creating a scroller.

Both pages now start with a doctype and clip rather than hide. If a page renders and refuses to scroll, check which element is the scroller before checking anything else.

Note 0308 Sep 2026

The router was counting its load after the drop

Routed experts hold 1104M of Synth-2's 1.23B parameters. They were producing about 22% of the feed-forward output — a quarter of the model doing the work of a fifth of a layer. The first suspicion was the router's softmax. It was wrong.

Expert load was being tallied after tokens had already been dropped for capacity. An expert that was oversubscribed therefore looked perfectly healthy in the statistics: everything above its capacity had vanished before the count. The balancing bias, which exists precisely to move work away from an overloaded expert, saw nothing to correct and never fired. The imbalance was self-concealing.

The fix is one line of accounting — count by demand, before capacity is applied — and the expert share now prints itself in every log line rather than needing a special run to measure.

The next session confirms or refutes this for free. A metric that is computed downstream of a lossy step is not measuring what its name says.

Note 0207 Sep 2026

A lower learning rate buys a step, not a slope

Held-out loss over steps 37,500–82,000, in three regimes. The middle row is what a learning-rate drop looks like while it is still paying out; the bottom row is the same drop a few thousand steps later.

RegimeSteps measuredHeld-out loss per 1,000 steps
Previous learning rate12,450−0.0077
LR ÷ 3, first thousands3,550−0.0602
LR ÷ 3, thereafter16,000+0.0018

Seventeen samples across the last session: least-squares slope +0.00178 per thousand steps, residual noise 0.017, t = 2.08. There is no descent left in it. Lowering the rate moves the level at which the loss settles and does not restore a slope — which means a third cut would buy a third step down and then flatten again.

The learning-rate probe has been retired. It measures the step, not the slope, so it will answer lower it every time it is asked. Two explanations survive: the sparse branch running at a fraction of its size (see note 03), and gradient noise at a 65K-token batch under an optimiser whose step length barely depends on the gradient's magnitude.

Note 0107 Sep 2026

Model-flops utilisation is 5.9%, and the obvious suspects are cleared

Attention, cleared by two measurements. Blocks of 256 beat all three alternatives that were tried, and switching the blocking off entirely makes the step 7.3% slower — so the usual fix reached for at this symptom, splash attention, would win nothing here.

Expert capacity, cleared. Raising it recovers about 14% of dropped tokens and costs roughly 60% more in the mixture-of-experts layer. Not a trade worth making at this size.

One candidate is left: the Newton–Schulz iteration inside the optimiser, which runs orthogonalisation on every update. The sweep that settles it is staged and takes about half an hour of TPU time.

Three measurements, three suspects removed, no code changed. Most of the value of a profiling session is the branches it lets you stop thinking about.

Nothing here is a result: the run is a third of the way through a thirteen-billion-token budget and no acceptance criterion has been measured yet. The live state of the run is on the Synth-2 page, read from the log rather than typed in. If a number here looks wrong, saying so is worth more than a compliment — ssanelabs@gmail.com.