Motion Style Slider: Endpoint-Supervised Continuous Style Control for Human Motion Diffusion
Abstract
Existing human motion diffusion methods provide strong motion generation quality [35], and recent style transfer models can inject target style cues [11,33], but fine-grained continuous control of style intensity remains underexplored. In production, style intensity is subjective across artists and directors, so the practical requirement is not a universal absolute unit (e.g., globally correct “2×”), but a reliable monotonic control axis. We propose Motion Style Slider, a motion-to-motion style transfer framework for endpoint-supervised continuous control. Given a content motion and a style motion, we construct a style direction in a learned motion-style embedding space and condition diffusion generation with a scalar intensity α. The training objective combines diffusion denoising with latent intensity regularization to encourage smooth and monotonic style scaling without requiring intermediate-intensity groundtruth motions. Our framework is compatible with pretrained motion diffusion backbones and supports heterogeneous style datasets, including the multi-actor style motion dataset [20]. To test out-of-range usability, we additionally introduce a small real-capture over-reaction extension and evaluate large-α behavior against these unseen targets. Experiments measure controllability, interpolation/extrapolation behavior, content preservation, and motion realism, with ablations on direction construction and loss design.