AI Safety
Where interpretability meets reliability.
My current approach to AI safety emphasizes interpretability and formal evaluation. Both are measurement practices: limits in what they detect should place corresponding limits on the guarantees we draw from them.
One hypothesis I want to test is whether sparse-autoencoder features remain stable across the boundary between clear and ambiguous inputs. Instability could warn that an interpretability profile does not transfer into that region. That relationship is plausible, but not yet established by my work.
The broader methodological commitment is to report the instrument, validate against independent references where possible, and state what a result does not show. The limits-to-formalization thread is a conceptual companion to this position, not empirical evidence for it.