CORE-V: Chain-Of-thought REasoning for Image Editing with Visual Interaction
Abstract
Existing methods for instruction-guided image editing are usually confined to text-based CoT reasoning for semantic understanding and reasoning, and lack of visual awareness necessary for fine-grained synthesis. In this work, we propose CORE-V, a novel Visual CoT framework to explicitly perform visually-interactive reasoning process for controllable image editing. Specifically, CORE-V refines global semantics into detailed entities and generates visual imagination images based on attribute description in the instructions. These visual imagination images, together with spatial conditions of input images (i.e., segmentation masks, depth maps and edge maps) are applied within the reasoning process. To achieve visual reasoning training, we build the CORE-Edit-800K dataset that contains both reference images and spatial controls aligned with reasoning pipelines. Furthermore, we introduce CORE-Bench to assess directional editing similarity. Comprehensive experiments on EmuEdit, Reason-Edit and CORE-Bench demonstrate superior performance of CORE-V for visual-semantic alignment and structural integrity.