SafeSAE-VLA: Interpreting OpenVLA Progress Dynamics with Sparse Feature Analysis
Abstract
Understanding whether modern vision-language-action poli-cies internally represent task progress is central to interpretable robotcontrol. This paper presents a progress-based SafeSAE-VLA analysis ofOpenVLA on 750 LIBERO episodes, replacing a degenerate binary targetwith a suite-normalized high-vs.-low split over relative geometric progressand testing layer-20 sparse-autoencoder (SAE) features (dsae = 16,384)with non-parametric differential analysis and false-discovery-rate (FDR)control. The recovered signal is strong and sparse. Of 1,881 active fea-tures, 1,117 are significant (FDR < 0.05), a linear probe reaches 0.918AUROC (area under the ROC curve), and top-20 features alone retain0.894 AUROC. Performance is stable across suites (up to 0.985 on goaland 0.947 on object), while motion-only controls are markedly weaker(0.572–0.711), indicating separability is not explained by gross move-ment magnitude. Dense raw-activation probes are equal or stronger thanSAE readouts (e.g., raw LR 0.975 vs. full-SAE 0.947 under a pooledanalysis with bootstrap confidence intervals). We therefore position thecontribution not as predictive dominance but as showing that nearly allof the progress signal survives in a compact set of named, inspectable,intervention-compatible directions. Layer sweeps, split-robustness checks,and a success-labeled audit (in which the progress-tuned top-20 correctlyscores an at-chance 0.471 while the broader SAE basis reaches 0.968)support this scoping, and direct on-OpenVLA feature-setting interven-tions produce specific, task-local behavioral shifts rather than generalclosed-loop repair.