InsertAnywhere: Geometrically Grounded and Optics-Aware Video Object Insertion
Abstract
Recent advances in diffusion models have enabled impressivevideo editing capabilities, yet production-grade Video Object Insertion(VOI) remains challenging due to inadequate 4D scene understandingand a lack of proper optical interactions, such as shadows and reflec-tions. To address these limitations, we present InsertAnywhere, a com-prehensive VOI framework that achieves geometrically grounded objectplacement and optics-aware video synthesis. Our approach first lever-ages a 4D-aware mask generation module that allows users to anchor anobject’s 3D pose in a single frame. The framework automatically prop-agates this placement across the video, accurately handling local scenedynamics and occlusions. To synthesize realistic physical lighting inter-actions, we introduce Optics-Aware Representation Alignment, a novelstrategy that utilizes an extended mask to guide feature extraction, en-abling optical effects to seamlessly extend beyond the inserted object’sboundary. Finally, to overcome the lack of training data for such phenom-ena, we construct and open-source ROSE++, a specialized quadrupletdataset tailored for the supervised learning of optical effects. Extensiveexperiments demonstrate that InsertAnywhere produces geometricallyplausible and photometrically realistic insertions in complex real-worldscenarios, significantly outperforming existing research and commercialgenerative tools.