DriveWeaver: Point-Conditioned Video Inpainting for Controllable Vehicle Insertion in Autonomous Driving Simulation
Abstract
A pivotal step in autonomous driving simulation involves in-serting foreground vehicles with predefined trajectories into simulatedscenes. This process enhances scene diversity and facilitates the creationof various corner cases for testing and improving autonomous drivingmodels. However, existing methods often rely on pre-reconstructed 3Dassets, which frequently lead to lighting inconsistencies between the in-serted foreground and the background. Moreover, the reliance on limited,manually-curated 3D assets hinders large-scale deployment. To addressthese challenges, we propose DriveWeaver, a novel framework for con-trollable vehicle insertion in autonomous driving simulation. Specifically,for a masked target insertion area, DriveWeaver performs video inpaint-ing conditioned on vehicle point clouds to generate high-quality, tempo-rally consistent vehicles. This video-inpainting-based approach ensuresseamless blending between the foreground and background, while thereadily available point cloud conditions enable superior generalization.To support long-term generation, we further design a global-to-local hier-archical inpainting strategy, ensuring the consistent identity and appear-ance of the inserted vehicles. Meanwhile, we extract explicit 3D Gaussianrepresentations of the inserted vehicles through an urban reconstructionpipeline to enable real-time rendering for autonomous driving simula-tion. Extensive experiments across diverse datasets demonstrate thatour method outperforms existing baselines in visual realism and geomet-ric consistency, providing a robust tool for scalable autonomous drivingscene augmentation.