LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
Abstract
Recent 3D human motion generation models demonstrateremarkable reconstruction accuracy yet struggle to generalize beyondtraining distributions. This limitation arises partly from the use of pre-cise 3D supervision, which encourages models to fit fixed coordinate pat-terns instead of learning the essential 3D structure and motion–semanticcues required for robust generalization. To overcome this limitation, wepropose LaxMotion, a framework that synthesizes realistic 3D motionswithout direct 3D pose supervision. Instead of regressing toward exactcoordinates, LaxMotion learns 3D motion as a consistent explanationof global trajectories and monocular 2D kinematic cues. We introducea structured motion factorization together with a reformulated trainingparadigm under relaxed observability. This design is further supported byrelaxed regularization objectives that enforce view-consistent alignment,orientation coherence, and structural stability. Under this relaxed super-vision paradigm, LaxMotion generates diverse, temporally coherent, andsemantically aligned 3D motions, achieving performance comparable toor surpassing fully 3D-supervised methods. These results indicate thatshifting supervision from exact coordinate matching to structural consis-tency promotes stronger reasoning and improved generalization, offeringa scalable and data-efficient paradigm for 3D motion generation.