MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation
Abstract
Recent advances in Diffusion Transformers (DiTs) have en-abled high-quality joint audio-video generation, producing videos withsynchronized audio within a single model. However, existing control-lable generation frameworks are typically restricted to video-only con-trol. This restricts comprehensive controllability and often leads to sub-optimal cross-modal alignment. To bridge this gap, we present MM-Control, which enables users to perform Multi-Modal Control in jointaudio-video generation. MMControl introduces a dual-stream conditionalinjection mechanism. It incorporates both visual and acoustic controlsignals—including reference images, reference audio, depth maps, andpose sequences—into a joint generation process. These conditions are in-jected through bypass branches into a joint audio-video Diffusion Trans-former, enabling the model to simultaneously generate identity-consistentvideo and timbre-consistent audio under structural constraints. Further-more, we introduce modality-specific guidance scaling, which allows usersto independently and dynamically adjust the influence strength of eachvisual and acoustic condition at inference time. Extensive experimentsdemonstrate that MMControl achieves fine-grained, composable controlover character identity, voice timbre, body pose, and depth-guided scenestructure in joint audio-video generation.