Co-Steer: Cross-Modal Collaborative Steering for Jailbreaking MLLMs
Abstract
Multimodal Large Language Models (MLLMs) remain vul-nerable to jailbreak attacks exploiting cross-modal interactions. Existingcross-modal attacks rely only on cross-entropy optimization, causingvisual and textual perturbations to act independently and push modelrepresentations toward misaligned vulnerability subregions, limiting theircombined effectiveness. To address this, we propose Co-Steer, a universalattack framework built on a key insight: visual and textual jailbreaks,despite appearing misaligned, are projections of a unified vulnerabilitydirection in the representation space. Co-Steer first extracts a universalsteering vector by aggregating representation shifts from successful vi-sual and textual jailbreaks via SVD. It then optimizes a single universalperturbation pair across queries using a steering alignment loss that coor-dinates both perturbations toward the extracted direction, transformingmisalignment into coordination. Analyzing such attack mechanisms helpsexpose gaps in current multimodal alignment and informs more robustdefense design. Experiments show our method achieves state-of-the-artattack success rates on MM-SafetyBench, HarmBench, and AdvBenchacross InternVL2-8B, Qwen2-VL-7B, and MiniGPT-4-13B, with notableimprovements in transferability to commercial models. Warning: Thispaper might contain harmful content.