Composing Driving Worlds through Disentangled Control for Adversarial Scenario Generation
Abstract
A major challenge in autonomous driving is the “long tail” ofsafety-critical edge cases, which often emerge from unusual combinationsof common traffic elements. Synthesizing these scenarios is crucial, yetcurrent controllable generative models provide incomplete or entangledguidance, preventing the independent manipulation of scene structure,object identity, and ego actions. We introduce CompoSIA, a composi-tional driving video simulator that disentangles these traffic factors, en-abling fine-grained control over diverse adversarial driving scenarios. Tosupport controllable identity replacement of scene elements, we proposea noise-level identity injection, allowing pose-agnostic identity generationacross diverse element poses, all from a single reference image. Further-more, a hierarchical dual-branch action control mechanism is introducedto improve action controllability. Such disentangled control enables ad-versarial scenario synthesis—systematically combining safe elements intodangerous configurations that entangled generators cannot produce. Ex-tensive comparisons demonstrate superior controllable generation qualityover state-of-the-art baselines, with a 17% improvement in FVD for iden-tity editing and reductions of 30% and 47% in rotation and translationerrors for action control. Furthermore, downstream stress-testing revealssubstantial planner failures: across editing modalities, the average colli-sion rate of 3s increases by 173%.