Quiet Neurons

a 397,000-parameter Dragon Hatchling (BDH, arXiv:2509.26507), running in this tab

starting The full case, with every limitation →

It goes quiet on what it just learned,
not on what is easy to guess.

one pass · 77 letters · layer 2  
How many neurons fire 
firing = a strictly non-zero value after the rectifier, so there is no threshold we chose · warm-up = the same 13 letters in every training sequence, so it lives in the weights · repeats = a word it met seconds ago, so it can only be in the context
How surprised it is 

 

Computed in this tab from the shipped weights; nothing here is a recording. One word is one sequence and single words scatter, which is what the eight-word button is for. This is the signature, not its cause, on a toy model and a synthetic task. The full page has the architecture, the provenance ladder, a dense Transformer trained on the same task that shows none of this, and every limitation we found. One-page summary · check it against PyTorch · source