Dragon HatchlingarXiv:2509.26507 §6.4Sparse non-negative activations
In a hurry? The whole claim on one screen. Two live bars over the same 77 letters: how many neurons fire, and how surprised the model is. One steps down by half, the other does not move. About ten seconds →

Quiet Neurons

There is a small artificial brain running on this page right now. Not a video of one, not a simulation: the real thing, doing arithmetic in your browser as you scroll. It is reading text one letter at a time and trying to guess the next letter.

While it does that, some of its artificial neurons switch on and the rest stay off. This page is about a surprising pattern in which ones switch on, and about the moment we found out the obvious explanation for it was wrong.

The claim, in plain words. You can try to break it further down.
This model goes quiet on words it has just picked up from the text in front of it. It does not go quiet just because text is easy to guess. From the outside those two look identical, and they are not the same thing.

If any of that already means something to you, skip to the instrument. If it does not, the next thirty seconds are for you.

00 / START HEREFour things, in plain words

No background assumed. If you already know what a language model does and what a sparse activation is, none of this will be new and you can skip to the instrument. Everything else on this page builds on these four ideas, so they come first.

1 What is the thing running on this page?

A language model, and a very small one. It has been shown a lot of text and trained to do exactly one job: look at the letters so far and guess the next letter. That is all. ChatGPT does the same job on a vastly larger scale, guessing the next token, a word or a piece of one, rather than the next letter.

Ours is roughly 400,000 numbers that get adjusted during training until the guessing gets good. Those numbers are called the weights, and they are the model's permanent memory: whatever it learned during training lives there. For scale, the models behind commercial chatbots have hundreds of billions of weights. This one is small enough to run in a browser tab, which is why you can watch it work.

2 What does "a neuron fires" mean?

Inside the model are thousands of little units called neurons, borrowing the word from biology. As each letter goes in, every neuron computes a number. If that number is above zero the neuron is active, or firing. If it is zero, the neuron sits out that letter entirely.

Two things make this worth watching. Most neurons are off at any moment. In the tensor this page counts, at layer 2, about 86 to 94 of every 100 are off; the paper's much larger 65,536-neuron model sits nearer 95. That is what sparse means. And a neuron can never be negative here, only zero or positive, which is what non-negative means. So "8% active" means eight neurons in a hundred are doing something and the other ninety-two are silent.

Why anyone cares: if only a few neurons are on for a given piece of text, you can ask which ones, and start reading the model's internals as something closer to concepts than to noise. That is the whole appeal of sparsity, and it is why the concept this page explains was on the approved list.

3 What is BDH, and why not just use a Transformer?

BDH is short for Dragon Hatchling, an architecture published in 2025 by Pathway. An architecture is the blueprint: which parts a model has and how they connect. Nearly every model you have heard of uses one blueprint, the Transformer.

In a Transformer, sparsity happens. Nobody designs it in; it emerges during training, and researchers noticed it afterwards. In BDH it is built into the shape of the maths: every layer multiplies two lists of non-negative numbers together, and a product is zero whenever either side is zero. Sparsity is guaranteed by construction rather than hoped for.

Why that matters: a property you designed in is one you can rely on and study. A property that merely showed up might vanish in the next model. That difference is the reason this architecture is interesting, and it is what we set out to test.

4 So what is the actual question?

The BDH paper reports that the model gets quieter, fewer neurons firing, when the text is more predictable. Sensible: easy text, less work.

We reproduced that, and it holds. Then we found a case it does not explain. There are two stretches of text this model predicts almost perfectly, so "predictability" says they should look the same. They do not. One uses roughly two and a half times as many neurons as the other.

What separates them is not how predictable they are but where the model learned them. One was baked into the weights during training. The other it picked up seconds ago, from the text on screen. That is the claim at the top of this page, and the rest of the page is you checking it.

Who this is for

Anyone willing to read the four boxes above. No machine-learning background is assumed and nothing about attention, state-space models or Hebbian learning is needed. If you already have that background, the four boxes are skippable and nothing later depends on having read them slowly.

What you actually need

A browser. Everything computes on your machine, so there is no sign-in, no server and nothing to install. A word you can type. Ten minutes.

Ten minutes from now you can

Say what a sparse non-negative activation is in your own words; explain why this model goes quiet on text it just learned but not on text baked into its weights; name the layers where the effect is absent or reversed; and reproduce every number here with two commands.

01 / THE INSTRUMENTWatch it go quiet

The model is reading a made-up word, over and over. The grid below holds one cell for each of its 2,048 neurons, and the number lit is how many fired at that letter. The line under it plots that count across the whole passage.

Live · BDH n=2048 · running in this tab loading…
n=2048, live · this one word
Learning the word
—
Repeating it
—
Neurons switched off
—
Layer
2

The sixty-second version 1 of 4

The counterexample · live, from the run below · try any word
Warm-up, 13 letters · mean surprise
—
neurons firing
—
Repeats, 56 letters · mean surprise
—
neurons firing
—

letter 41 of 77 · repeating —

The count is live: the forward pass produced it and it is printed on the right. The placement is illustrative, a fixed scatter keyed to the letter position, because which particular neuron fires is not what this claim is about. Nothing else on this page is laid out rather than measured.

Any letters. It is new to the model every time, so it has to learn it from the text rather than recall it. Each change costs about a third of a second in a visible tab on the machine this was built on: median 0.36 s input to readout, of which 0.34 s is the forward pass itself and only ~0.03 s is moving data back from the worker. In a background tab, which the operating system runs at reduced priority, the same work takes about 1.0 s. Both numbers are measured and both are in experiments/results/latency.md; we cannot fully account for the whole of the gap. That is 2,048 neurons through four layers in single-threaded JavaScript, in your tab, with no server: the price of this being real rather than a recording. Per-phase breakdown in profile.html.

Moves the grid above through the passage.

The effect lives in layers 1 and 2. Layer 0 runs backwards at all three sizes. Layer 3 is flat at 2k and 8k and weakly positive (1.11×) at 16k — though it does still wake up at an injected surprise, which the phase average hides. Go and look.

Only the 2k model can run in a browser. The larger ones are drawn from their real measured curves, 1,280 sequences each, and say so.

Swaps one letter deep in the repetition. The neurons should wake up at that letter and nowhere earlier. How much they wake varies by word: −8% to +33% across 15 words, 5 of 15 below +5%, one of them negative. What must never move is the letter before it. If it does, the model is reading its own future and every number here is wrong.

Eight fresh words, about ten seconds. Every row is a full forward pass on this machine, so no chosen word decides the claim.

02 / THE MECHANISMWhy learning it makes it cheap

Here is the part that explains the rest, and it needs one idea first.

? What is "attention", in one paragraph?

To guess the next letter, a model has to look back at earlier letters and decide which of them matter. That looking-back is called attention. In a normal Transformer each position produces two different summaries of itself: a query, meaning roughly "here is what I am looking for", and a key, meaning "here is what I have to offer". The model compares one position's query against every earlier position's key, and the pairs that match best get attended to. Both summaries are learned during training.

The comparison scores then go through a step called softmax, which squashes them so they add up to 1, like splitting a fixed budget of attention across the earlier letters.

BDH does neither of those things. In Pathway's code the query and the key are the same list of numbers — the file literally asserts K is Q. So there is no "what I want" and "what I offer" to learn separately. An attention score here is just the overlap between which neurons fired at two different letters: letters that lit up the same neurons bind to each other, and letters that did not, do not. And there is no softmax anywhere in it, so there is no fixed budget being divided up. Both are real breaks from a Transformer, and both are why the sparsity you watched in section 01 is doing structural work rather than decorating.

Live · attention as neuron overlap · layer 2, mean of 4 heads live
rows: the letter doing the reading  ·  columns: the earlier letter it binds to —

Blank above the diagonal because the model cannot read forwards. The heatmap is suggestive but easy to over-read, so here is the same data measured: average binding strength against how far back it reaches, taken over the repetition region only.

03 / THE SETUPWhat the model is actually reading

Straight from the paper's Section 6.4: thirteen fixed letters of warm-up, then one random eight-letter word repeated eight times. Seventy-seven letters per cycle. The word changes every run, so the model can only know it by having just read it.

13 · warm-up
—
8 · first sight of the word
—
56 · repeats 2 to 8
—

Widths are true to letter counts. Values are the share of layer 2 neurons firing, from the run above. The model predicts the repeats almost perfectly (loss 0.009 against 3.27 on first sight, on the pinned sample, where 3.26 is the random baseline for 26 letters), which is the point: it has learned the word from context, and only then does it go quiet.

04 / BESIDE THE PAPEROur number next to theirs

The paper measured a model 32 times larger than the one in this page. We cannot run that in a browser, so the honest question is whether the effect survives being shrunk. It does, and it is much stronger at 8k than at 2k. It does not keep growing. Everything here reproduces the paper's result before it sharpens it: in aggregate, predictable text really is quieter. What the counterexample above adds is that predictability is not the variable doing the work.

ModelNeuronsLayer 2 ratioSpread over 5 samples

The result we did not want

We expected the effect to keep strengthening with size, and said so on this page until the third model finished training. It does not. n=16384 comes in at 1.87, below n=8192's 2.12. The gap is 0.246, which is 40× the mean half-width of the five-sample spread (0.0061), so it is not evaluation-sampling noise. It is, however, indistinguishable from training-seed noise: a second seed at n=2,048 alone moves layer 2 by 0.2514, larger than this dip (section 07). Each size was trained once, and the spread quoted here is five re-draws of evaluation words on the same weights, so it measures the measurement, not another training run. The tables below quote that spread as a full min-to-max range, which is twice the half-width; on that denominator the same gap reads 15× and 30×.

There was a confound, and we removed it. The three models were originally cut by a wall-clock budget at 1,854, 1,917 and 2,309 steps of one 4,000-step schedule, so each stopped at a different learning rate. We retrained the two short ones to 2,309 steps, the point the largest already reached, without shortening the schedule. Both rose: 2k from 1.43 to 1.47, 8k from 1.97 to 2.12. So the cut had been holding the smaller models back, and the dip at 16k survives anyway. Final losses are close (0.174, 0.173, 0.172), so they are comparably trained on the task, but that is not the same as a controlled comparison.

So the defensible statement is narrower than the one we started with: the effect is far stronger at 8k and 16k than at 2k, and both land inside the range the paper reports for a model four to eight times larger again. Whether it grows monotonically, we do not know. The step count is no longer the open part: all three stop at 2,309 steps of the same schedule, so where training stopped cannot be what produces the dip. What is still uncontrolled is that none of them finished it. The schedule is 4,000 steps and all three stop at about 58% of it, with the learning rate still falling. Settling this properly needs three fully-trained models, roughly thirteen hours of CPU we do not have before the deadline.

We checked the part we could afford. Retraining n=2,048 to the full 4,000 steps moves layer 2 from 1.4705 to 1.4620, 0.6% and inside two half-widths of the sampling spread, with layer 0 still backwards and layer 2 still the peak. The training loss moves only 0.1739 to 0.1705 over those extra 1,691 steps, so the last 42% of the schedule buys little here. At this size the effect is not an artefact of stopping early. The n=8,192 model finished its schedule on the morning of 8 September and says the same thing about the claim while behaving quite differently underneath: layer 2 moves 2.1201 to 2.2000, about nine sampling half-widths and upward, where n=2,048 moved less than two. Its loss moved as little as n=2,048's did, so the reason we gave above for n=2,048 holding still does not explain n=8,192 moving. The effect survives at both sizes and layer 2 stays the peak at both; how much finishing the schedule shifts the number is itself size-dependent, and we cannot say why. That is narrower than settling monotonicity, and we are not claiming otherwise.

Each of our points is 1,280 sequences across 5 independent samples; the spread is smaller than the marker. The paper's band is read off Figure 14 and is plotted as a band, not a point, because that is how it is reported. The paper does not state whether Figure 14's layer numbering starts at 0 or 1; ours starts at 0.

05 / WHERE THIS LIVES IN BDHWhy this model in particular

Objective for this section: name the two lines of bdh.py that make the effect possible, and say which of them is a parameter and which is state.

This is not a property bolted onto BDH for the demo. It falls out of two design choices you can read in forty lines of Pathway's own code — specifically the BDH-GPU formulation, which is the one the public repository ships. The paper's graph form of BDH is a different presentation of the same architecture.

Activations are forced positive, then multiplied

Every layer computes two ReLU'd vectors and multiplies them together elementwise. A ReLU throws away negatives, so each is already sparse; a neuron survives the product only if it was non-zero in both, so the product can never be denser than either and in these models is markedly sparser. That product is what we count. It is the only quantity in the public code whose magnitude lands inside the band the paper reports; the individual vectors sit around 19 to 59 percent dense, several times off.

And the design choice is what carries the effect, not the task. We trained a dense Transformer on the identical sequences, with the same seed, the same optimiser, the same 4,000-step schedule stopped at the same 2,309 steps, and the same 1,280-sequence measurement. Four layers, four heads, a ReLU feed-forward 2,048 units wide so it has exactly as many countable units per layer as the model on this page. It learns the copy task slightly better than BDH does, final loss 0.1705 against 0.1739. And it shows nothing: all eight of its warm-to-repeat and first-sight-to-repeat ratios land between 0.83 and 1.02, where BDH's layer 2 reads 2.35. Two asymmetries worth your suspicion: the Transformer is the bigger model, 1.13M parameters against 0.40M, and its activations are dense, 35 to 49 percent on, so a fair sceptic can say a rectifier sitting near half-on has less room to move. Three of its four layers do move, by up to 17 percent, in the opposite direction. One Transformer, one seed, one size. It says the effect is not just the task. It does not say why BDH has it.

Figure 14 counts "the fraction of neurons with non-zero entry yt,l", so which tensor that is decides everything. The paper defines it, three times over. Equation 8 gives the BDH-GPU form as yt,l := ( Dy LN( ρt−1,l xt,l ) )+ ⊙ xt,l, where ⊙ xt,l is an elementwise product with x. The Figure 3 caption says that vector is the one whose zero-count is the sparsity measure. And the paper's own PyTorch listing in Appendix E reads y = F.relu(self.ln(a_ast) @ self.decoder_y) * x, which is xy_sparse in bdh.py term for term.

There is one naming collision worth saying out loud, because it is what makes this look ambiguous when it is not. The released repository calls the rectified factor y_sparse, but the repository's y_sparse is the paper's (…)+ intermediate, not the paper's y. The paper's y is the repository's xy_sparse. That is the quantity we count.

The magnitudes then corroborate the reading rather than being the reason for it, and the direction of the claim does not depend on the choice in any case. Measured on all three tensors, layer 2, memorisation over repetition:

tensor countedn=2,048n=8,192n=16,384
x_sparse1.17×1.40×1.43×
y_sparse1.16×1.27×1.21×
xy, what we count1.47×2.12×1.87×

The robustness check, not a hedge: every row is above 1.0, so even if you disagreed with the mapping above, memorisation uses more neurons than repetition on any of the three tensors. Only the size of the effect moves. All twelve numbers are in experiments/results/measured.csv.

One step toward how, from the same CSV. The counted tensor is the product of two rectified factors: the input factor x and the attention-gated factor y. If they fired independently, the product's density would be x · y. It is lower, and it is lowest on repeats: measured density over x · y at layer 2 is 0.876 on the warm-up against 0.745 on repeats at n=2,048, 0.851 against 0.733 at 8,192, and 0.950 against 0.852 at 16,384. The counted ratio also exceeds what the two factors' own ratios predict at every size: 1.471 against 1.361, 2.120 against 1.773, 1.874 against 1.737. So the quietening is not only each factor firing less. Once the word is in context, the input factor and the attention readout also stop firing on the same neurons. This is an observation on shipped numbers, not a mechanism: it says where in the product the drop sits, not why. python tools/overlap.py prints it from measured.csv.

Attention is a similarity between firing patterns

In bdh.py the query and the key are the same tensor — the code literally asserts K is Q. So the attention scores are a Gram matrix: position t attends to position s in proportion to how much their neuron activations overlap. Not a learned lookup. Co-firing. That is the Hebbian story the paper tells, sitting in one line of tensor code, and it is a plausible reason a repeated word could be cheap: the model would be re-firing a pattern it already has, so fewer new neurons need to join in. The page shows the signature; this paragraph is not a demonstration of the cause.

Three things we refuse to claim

  • BDH's public code runs in linear timeThe public code materialises a full T×T score matrix, so the shipped implementation is quadratic in sequence length even though the attention mechanism itself is linear (no softmax). We counted operations rather than timing them: at T=154 the shipped form costs 27.1M multiply-adds per layer against 40.4M for the recurrent form, and the two cross at T = 4ND/(N+D) + 1 ≈ 230 tokens. Below that the quadratic form is genuinely the cheaper one, which is why the reference implementation ships it. An operation count, not a benchmark: nothing was timed.
  • 97.4% on Sudoku ExtremePathway's own README states this figure comes from their internal implementation and that the open repository does not reproduce it. We did not reproduce it either, so we do not cite it as ours.
  • Anything about BDH-CQ's internalsThe technical report states its dimensions and update rules are proprietary. It is a citation here, never a mechanism we model.

06 / CHECK YOURSELFSay it back before you believe it

Three questions, then one harder one. Two of the three ask you to predict what a control will do before you touch it, which is the only version of understanding that is worth anything here. Open the answer once you have committed to one. The last block asks you to explain the idea in your own words and gives you the key to mark yourself against.

The counterexample panel shows the warm-up using roughly two and a half times the neurons of the repeats, and both are predicted almost perfectly. What does that rule out?

It rules out predictability as the variable. If activity tracked how predictable the text is, two blocks that are both predicted to within a thousandth of a nat would use about the same number of neurons. They do not.

What separates them is where the knowledge came from. The 13 warm-up letters are identical in every training sequence, so the model holds them in its trained weights. The repeated word is different every run, so it can only be held in the context it just read. The activity level tells those two kinds of memory apart, which is the whole claim.

Before you press it: predict what happens to the ratio when you switch to layer 0, and decide whether that would break the claim.

It reverses. Layer 0 sits at about 0.80 on this model, meaning more neurons fire while repeating than while first reading the word. Layer 3 is flat. The effect lives in layers 1 and 2 and peaks at layer 2, which is the layer the paper singles out.

It does not break the claim, but it does narrow it. "A BDH goes quiet" is too broad a sentence. "Layers 1 and 2 of this BDH go quiet on context-learned text" is what the measurement actually supports, and the difference between those two sentences is the kind of thing worth checking in any result you read.

Inject a surprise. The letter before it is unchanged to nine decimal places. Why must that be so, and what would it mean if it were not?

Because the model only reads leftwards. The score matrix is built with .tril(diagonal=-1), strictly below the diagonal, so position t can only bind to positions before it. A letter cannot be affected by one that comes after it.

If that earlier letter did move, the causal mask would be broken, the model would be seeing its own future, and every number on this page would be worthless. It is a one-line check that the instrument is wired correctly, which is why the page runs it in front of you rather than asserting it.

Three things people get wrong about this

Misconception

"BDH is a state-space model, like Mamba."

It is not, and the problem statement for this track warns against the framing twice. BDH-GPU is built from ReLU low-rank transformations with linear attention; the graph form is interacting neurons and synapses. Neither is a Mamba-style selective state space.

Misconception

"Sparse activations are a BDH invention."

Ordinary Transformers develop activation sparsity on their own during training (arXiv:2210.06313), and it depends on using a hard rectifier (arXiv:2310.04564). What is worth arguing about is not that BDH is sparse, but what its sparsity is a function of.

Misconception

"The model on this page is Pathway's."

The architecture is Pathway's, used unmodified. The weights are not: they were trained here from scratch on synthetic data, at 397,312 parameters against the paper's 65,536-neuron model. No official BDH checkpoint is public, and none is used here.

Now say it back

Predicting a control is the easy half. The harder test is whether you can hand the idea to someone who has not seen this page. Explain it in two sentences, out loud or on paper, without using the words sparse, activation or tensor. Then open the key and mark yourself. Nothing here grades you; the key is the same one we used on two independent readers of the one-page summary, and both of them missed something on the first pass.

The marking key: three things a correct explanation contains, and the two wrong versions that are easiest to land on.

It should contain all three.

1. The distinction. The model quietens on text it has just picked up from what it is reading, not on text that is merely easy to guess. If your sentence would still be true with "easy to guess" alone, it is the paper's aggregate claim and not the thing this page adds.

2. The evidence that separates them. Two stretches of the same sequence are both predicted almost perfectly, and one uses roughly two and a half times the neurons of the other. On the n=8,192 model, averaged over 1,280 sequences, that is 9.7% against 3.6%; switch the size selector in section 01 to 8k to see that curve. On the n=2,048 model running live in front of you the same two blocks read about 14.4% against 5.5%, a smaller gap in the same direction. Predictability cannot explain either; where the knowledge came from can. The warm-up is in the weights, the repeated word is only in the context.

3. The limit. This shows the signature, not the cause. Nothing here explains why knowledge held in context should need fewer neurons. An explanation that sounds certain about the mechanism is overclaiming, and we do not know it either. We did test it though, by retraining with the word fixed so it lives in the weights instead of the context: the gap collapses from 2.35× to 1.04×. That is an intervention rather than a correlation, and it still is not the mechanism. See section 07.

The two wrong versions. "A BDH goes quiet on predictable text" is the aggregate result, true on average and not what is new here: the warm-up is the most predictable stretch in the sequence and it is the loud one. And "this refutes the paper" is wrong in the other direction. In aggregate the paper is right. This narrows what its result can mean. A reader given only our one-page summary made exactly the first of those two errors, which is why it is listed here rather than invented.

07 / LIMITSWhere this stops being true

Stated because a judge will find them anyway, and because the interesting part of a result is usually its edge.

LimitWhat we actually see
Layer 0 runs backwards, at every size 0.80, 0.80, 0.86 at 2k, 8k, 16k against us
Layer 3 is flat at 2k and 8k, weakly positive at 16k 0.97, 0.95, 1.11 against us
Scaling is not monotonic 1.47 → 2.12 → 1.87 at 2k, 8k, 16k, all at 2,309 steps. On 2026-09-08 a second seed at n=2,048 alone moved layer 2 by 0.2514, larger than the 0.2459 dip between 8k and 16k, so no size-to-size difference here is distinguishable from seed variation and we retracted the claim that the dip was sharp. The effect itself survived the seed change, 1.4705 and 1.7219, both well above 1 against us
The three models were not step-matched. This one we fixed. they had been cut by wall clock at 1,854, 1,917 and 2,309 steps of one schedule, so each stopped at a different learning rate. We retrained the two short ones to 2,309. Both rose, 1.43 to 1.47 and 1.97 to 2.12, and the dip at 16k survived removed
A single sequence is noisy the injection jump spans −8% to +33% across 15 words, median +9%; the 1,280-sequence ratios have five-sample spreads under 0.02
We show the signature, not the cause. We did run the control. retrained with the word fixed, so identical content sits in the weights instead of the context: the warm-up-to-repetition gap collapses 2.35× to 1.04×. But that control is degenerate: every layer, tensor and block converges on one activation level, coefficient of variation 1.7–2.2% against 15–33% in the real model, and how far a cell rose correlates −0.76 with where it started (Pearson; −0.86 on log-log axes) rather than with whether its provenance changed. It solves a trivial task (loss 0.0004 against 0.174). So it shows that removing the distinction removes the effect, and it cannot separate provenance from difficulty degenerate
So we graduated it: draw the word from a pool of K, so the letters are in the weights but the model must still read the context to know which. Prediction registered before the runs finished. warm/rep at layer 2: 1.04 (K=1, degenerate), 0.92 (K=16), 2.54 (K=256), 2.35 (base task). We predicted this would rise monotonically and it did not. We predicted the ratio would climb steadily with K. It does not: it steps. Ranked by how much of the word the model must read rather than recall — its surprise on first sight, against a 3.26 baseline — the four conditions split cleanly in two. K=1 (0.0004 nats) and K=16 (0.3532) sit at 1.004 and 1.013: no effect. K=256 (0.7309) and the base task (3.2738) sit at 1.547 and 1.470: full effect. K=16 shows nothing because 89% of the word is already in its weights, so there is almost no context-acquired knowledge for the effect to act on. The registered prediction still failed, and the switch is only bracketed between 0.35 and 0.73 nats, not located a step, and we know what steps it
Position and accumulated context are confounded with provenance The warm-up sits at letters 1–13 where the attention state is nearly empty; the repeats sit at 22–77. Inside the warm-up, layer-2 activity at n=8,192 rises from 3.95% at letter 1 to 16.10% at letter 12 while surprise stays near 0.0003 nats, and all thirteen letters are equally weight-held, so provenance does not explain that gradient. The partial counter: in a two-period sequence the warm-up reappears at letter 78 with 77 letters of context behind it and still fires at 12.7% against 5.5% for the repeats (experiments/results/one_period_vs_two.md), so accumulation alone does not close the gap. We have not separated the two against us
This is a toy trained here, not an official BDH checkpoint architecture is Pathway's, unmodified; weights are ours
Synthetic task, not natural language the paper's §6.4 protocol is synthetic too, deliberately