Sen-Cap: Sensor-Flexible and Noise-Resilient Human Motion Capture via LiDAR-Camera Integration
Abstract
We propose Sen-Cap, a Sensor-Flexible and Noise-Resilient3D human motion Capture framework that integrates multi-modal datafrom LiDAR and camera. While multi-modal sensors provide richer infor-mation than single-modal sensors, existing approaches still sux001Ber fromtwo core challenges. First, multi-modal alignment/matching across ar-bitrarily deployed sensors is typically handled by explicit calibration,which propagates errors under changing viewpoints and in turn con-strains deployment to x001Cxed, highly overlapped layouts. Second, priormethods degrade under severe noise or partial sensor failures, which areSen-common in real-world environments. To address these challenges,Cap introduces a Unix001Ced Across-Sensor Motion Estimator that recon-structs local pose and shape in a human-centric space without calibra-tions between sensors, supporting a x001Dexible number of sensors, as well asa Noise-Resistant Trajectory Tracker that maintains robustness under se-vere point cloud noise through iterative rex001Cnement. These sensor-x001Dexibleand noise-resilient features make Sen-Cap more practical in real-worlddeployment. Notably, operating in real time, Sen-Cap achieves state-of-the-art performance on major metrics on Human-M3 and FreeMotion,as well as strong cross-domain performance on LiDARHuman26M andRELI11D. This combination of x001Dexibility and robustness opens new op-portunities for motion capture in real-world scenarios, e.g. sports ana-lytics, x001Celd robotics, and large-scale immersive environments.