Mechanistic interventions for explainable digital pathology uncovers adversarial vulnerabilities
Abstract
The deployment of deep learning in digital pathology holds transformative potential for cancer diagnosis, yet its clinical adoption is hindered by the opaque and vulnerable nature of black-box models. In this work, we introduce a mechanistic examination framework that combines sparse autoencoders and activation patching to dissect adversarial vulnerabilities in pathology models as several learned features may correspond to non-morphological artifacts rather than clinically meaningful patterns. By isolating the neural circuits responsible for these vulnerabilities, we demonstrate that models often rely on spurious, nongeneralizable cues, particularly in middle layers of Vision Transformers and later convolutional blocks of ResNets. To address this, we propose Mechanism-Informed Adversarial Training, a novel approach that targets vulnerable circuits with precision regularization, achieving state-of-theart robustness while preserving clean performance. Crucially, our framework aligns with rapidly emerging clinical and regulatory demands for alignment, accountability, and safety in AI-driven diagnostics. By providing a mathematically rigorous, interpretable, and actionable method to audit and fortify deep learning models, we seek to bridge the gap between AI research and clinical deployment, offering a pathway to trustworthy, regulation-ready AI in healthcare. Our findings not only inform the field of adversarial machine learning but also set a precedent for mechanistically grounded artificial intelligence in real-world medical applications.