(play from 0:00 on YouTube)
(play from 0:05 on YouTube)0:05 Beside “thinking about cats”, a fuzzy grey cat face is taped in: the 2012 “cat neuron”, the first famous cat found inside a network, which came out of a sparse autoencoder. Building high-level features using large scale unsupervised learning
(play from 0:05 on YouTube)0:05 “Thinking about cats” (and the cat it is still thinking about at the very end) is also the trait in Neel's subliminal-learning paper: “You love cats. You think about cats all the time.” Neel's activation-diffing paper’s agent read a model’s hidden love of cats too. Subliminal Learning Is Steering Vector Distillation
(play from 0:06 on YouTube)0:06 Beside the lens, a specimen tag on a string: “SPECIMEN No. 1, toy model, 1 layer”. The first specimen is a toy, as in Toy Models of Superposition, and one layer is where Neel’s grokking work lives: his paper fully reverse-engineered the algorithm a one-layer transformer learned for addition mod 113. Progress measures for grokking via mechanistic interpretability
(play from 0:06 on YouTube)0:06 The margin notes are headed “Observ. I.”, “Observ. II.”, as Hooke heads the chapters of Micrographia (1665). Zoom In ends by comparing interpretability to it. Zoom In
(play from 0:07 on YouTube)0:07 In its thought bubble, a cat; as it starts to grow, the front of a car crowds in beside the cat, then its thoughts tangle and cloud over. One InceptionV1 neuron, 4e:55, responds to cat faces and car fronts: Zoom In’s polysemantic neuron, the one verse 1’s lamp finds. Zoom In
(play from 0:08 on YouTube)0:08 As it starts to grow she hedges: a red caret inserts “mostly” into “it’s harmless”, the Hitchhiker’s Guide’s revised entry for Earth after fifteen years of field research. Mostly Harmless
(play from 0:10 on YouTube)0:10 Once the year reaches 2022 (when the library appeared), the brass lens is etched “TransformerLens”, small, just inside the rim. TransformerLens
(play from 0:10 on YouTube)0:10 On the big lens close-ups, under the TransformerLens etching, a fainter engraving scratched out: EasyTransformer, the library’s name in 2022 (“a transformer mechanistic interpretability library I’m writing called EasyTransformer”). Real-Time Research Recording






























![Margin card reading [don't][go] … [don't] → [go]](stills/refs/C1a-3.jpg)











































































































































































































































