Saber: Anchoring Semantics to Scale-Aware Kinetic Salience for Zero-Shot Skeleton Action Recognition
Abstract
Zero-Shot Skeleton Action Recognition (ZSAR) aims to rec-ognize unseen action categories by establishing a generalizable mappingbetween skeletal kinematics and semantic representations. However, ex-isting methods frequently struggle with two fundamental issues: (i) theloss of high-frequency dynamics caused by recursive feature aggrega-tion, and (ii) rigid semantic anchoring, which leaves cross-modal align-ments highly vulnerable to semantically irrelevant background noise.To overcome these limitations, we propose the Scale-Aware BipartiteEnergy-guided Registration (Saber) framework, a unified kinematics-driven paradigm for robust cross-modal alignment. Specifically: (i) Scale-Aware Bipartite Encoder actively reshapes the spatio-temporal topol-ogy and applies instance-adaptive filtering to effectively isolate high-frequency motion details, thereby preserving critical structural dynam-ics against “spectral smoothing”. (ii) A Kinetically Driven Probabilis-tic Alignment strategy, comprising energy-guided dynamic focusing anduncertainty-aware probabilistic anchoring, maps part-level representa-tions into variance-adaptive distributions. This mechanism dynamicallyresolves structural uncertainties and robustly anchors physical move-ments to explicit semantics. Extensive evaluations on various bench-mark datasets validate that Saber achieves state-of-the-art performance,demonstrating exceptional robustness and generalization.