VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On
Abstract
As virtual try-on (VTON) advances, a growing number ofreal-world scenarios have emerged, pushing beyond the ability of the ex-isting specialized VTON models. Meanwhile, universal multi-referenceimage editing models have progressed rapidly and exhibit strong gen-eralization in visual editing, suggesting a promising route toward moreflexible VTON systems. However, despite their strong capabilities, thestrengths and limitations of universal editors for VTON remain insuffi-ciently explored due to the lack of systematic evaluation benchmarks. Toaddress this gap, we introduce VTEdit-Bench, a comprehensive bench-mark designed to evaluate universal multi-reference image editing modelsacross various realistic VTON scenarios. VTEdit-Bench contains 24,220test image pairs spanning five representative VTON tasks with progres-sively increasing complexity, enabling systematic analysis of robustnessand generalization. We further propose VTEdit-QA, a reference-awareVLM-based evaluator that assesses VTON performance from three keyaspects: model consistency, cloth consistency, and overall image quality.Through this framework, we systematically evaluate 8 universal editingmodels and compare them with 7 specialized VTON models. Resultsshow that top universal editors are competitive on conventional tasksand generalize more stably to harder scenarios, but remain challengedby complex reference configurations, especially multi-cloth conditioning.