DisentangledTMR: Privacy-Preserving Skeleton Motion Retargeting via Factorized Transformers
Abstract
Skeleton-based motion data leak personally identifiable information through both static skeletal structure and dynamic motion patterns, enabling re-identification even without facial features. We present DisentangledTMR, a Transformer Motion Retargeting (TMR) architecture that achieves privacy through explicit architectural disentanglement. Two encoders with complementary inductive biases, temporal convolutions for action and spatial graph convolutions for identity, feed a factorized decoder that fuses their representations through separate crossattention streams and adaptive gating. A three-stage training curriculum progressively establishes disentanglement, reconstruction, and endto-end refinement, and a tunable partial-retargeting ratio trades privacy for compatibility with pre-trained downstream models. On three benchmarks, DisentangledTMR substantially reduces re-identification while preserving action recognition, outperforming single-encoder baselines.