FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control
Abstract
Generative video compositing, which involves inserting ex-ternal assets seamlessly into existing video sequences, is essential forcontent creation and visual effects. However, existing approaches suf-fer from a control-fidelity trade-off: they either hallucinate motion fromstatic images, failing to preserve the dynamics of pre-animated assets, orlack fine-grained spatial control for precise asset placement along user-defined trajectories. We propose FlexComposer, a unified framework thatstandardizes video compositing as a trajectory-guided conditional gen-eration task, enabling the seamless integration of both static images anddynamic footage. Our approach introduces three key designs: (1) a Uni-fied Canonical Foreground Representation that decouples an object’s in-trinsic motion from its global displacement, standardizing heterogeneousinputs into a stabilized, centered latent space; (2) a Spatial-Aware La-tent Injection strategy that exploits the translation equivariance of VAElatent spaces to transport canonical features onto target trajectories viaa parameter-free mechanism; and (3) a Hybrid Dataset and Synthetic-to-Real Curriculum that synergizes procedural simulation, real-world cine-matic footage, and generative data to implicitly learn physically plausibleillumination and shadow harmonization. This unified design handles di-verse inputs—from product photos to dynamic subjects—achieving high-fidelity motion control and environmental integration without the needfor explicit 3D reconstruction or auxiliary learnable adapters. Extensiveexperiments demonstrate that FlexComposer outperforms state-of-the-art methods in visual quality, temporal consistency, and trajectory ad-herence.