Mechanistic Interpretability
Inherited Circuits, Learned Semantics: How Fine-Tuning Creates Evasion Vulnerabilities Invisible to Standard Evaluation
A mechanistic interpretability study showing how security fine-tuning can specialize inherited model circuits into brittle indicator rules while preserving standard benchmark accuracy.
Applied Interpretability: Foundation-Sec-Instruct Goes Under the Microscope
Exploring mechanistic interpretability methods for understanding internal behavior of security-focused language models.