Agent-OBJ: Prompt-Driven 3D Adversaries for Multi-Modal Perception
Abstract
Modern autonomous driving systems increasingly rely ontightly coupled Camera–LiDAR fusion pipelines to achieve robust sceneunderstanding. While this multi-modal redundancy is designed to counterindependent single-sensor failures, it implicitly assumes that physicalobjects cannot simultaneously deceive both modalities. In this paper, weexpose the vulnerability of this assumption by proposing Agent-OBJ,a generative framework that synthesizes physically plausible 3D adver-saries via a prompt-driven agent. Unlike prior works that mainly focus onsensor-specific perturbations, our pipeline instantiates a physically plau-sible 3D adversary by generating a canonical pedestrian geometry from abase prompt and then modulating its shape via an additional semanticprompt, while controlling appearance with a pretrained multi-view per-sonalization model conditioned on multi-view images and a style promptto produce view-consistent appearance. To simultaneously compromiseboth 2D monocular and 3D fusion detectors, the agent dynamically op-timizes the generated adversary using confidence-score feedback fromboth detectors. Meanwhile, to ensure stealthiness, we regularize the gen-erated 3D point cloud by enforcing similarity to a pedestrian featurebank, aligning it with the distribution of natural objects and makingit visually and geometrically indistinguishable from benign instances.Furthermore, we incorporate a multi-view consistency constraint duringoptimization to promote cross-view robustness, ensuring the adversaryremains effective under diverse viewpoints in real-world driving scenes.Extensive experiments on nuScenes demonstrate high attack success ratesagainst both monocular and fusion detectors while preserving strongperceptual plausibility.