Sparse autoencoders make a bad alarm and a great microscope.
Sparse autoencoders are the interpretability advance of the last two years, and the instinct is to wire one in as a runtime monitor for a known concept. The benchmarks say not to: for detecting, probing, or steering a concept you can already name, SAEs lose to a logistic-regression baseline and flip under a one-token adversarial nudge. Their real strength is the opposite job — discovering concepts you did not know to look for.