praxagent / Technical record
Notes tagged Sparse-Autoencoders
Research notes collected under Sparse-Autoencoders, ordered by date.
A Linear Nudge, a Nonlinear Wake
We followed controlled hidden-state edits through the final 29 blocks of Llama 3.3 70B. The delivered edits were dose-linear and a fixed corpus-average Jacobian preserved that scaling, while the model's actual downstream …
Can a Jacobian Lens Detect SAE Steering?
A prospectively frozen Llama 3.3 70B experiment asks whether SAE steering leaves a detectable downstream fingerprint in Jacobian-lens space. A preregistered follow-up adds semantic hard negatives, same-subfamily …
How to Read an SAE Feature ID
A primer on sparse autoencoders: what a feature ID is, how labels get assigned, and why an activation map is not yet an explanation. A public deception/roleplay feature set is used as a worked example under Llama 3.3 70B …