DSH-Bench: A Difficulty- and Scenario-Aware Benchmark with Hierarchical Subject Taxonomy for Subject-Driven Text-to-Image Generation
Abstract
Significant progress has been achieved in subject-driven text-to-image (T2I) generation, which aims to synthesize new images depictingtarget subjects according to user instructions. However, evaluating thesemodels remains a significant challenge. Existing benchmarks exhibit criti-cal limitations: 1) insufficient diversity and comprehensiveness in subjectimages, 2) inadequate granularity in assessing model performance acrossdifferent subject difficulty levels and prompt scenarios, and 3) a pro-found lack of actionable insights and diagnostic guidance for subsequentmodel refinement. To address these limitations, we propose DSH-Bench,a comprehensive benchmark that enables systematic multi-perspectiveanalysis of subject-driven T2I models through four principal innovations:1) a hierarchical taxonomy sampling mechanism ensuring comprehensivesubject representation across 58 fine-grained categories, 2) an innovativeclassification scheme categorizing both subject difficulty level and promptscenario for granular capability assessment, 3) a novel Subject IdentityConsistency Score (SICS) metric demonstrating a 9.4% higher correlationwith human evaluation compared to existing measures in quantifyingsubject preservation, and 4) a comprehensive set of diagnostic insightsderived from the benchmark, offering critical guidance for optimizing fu-ture model training paradigms and data construction strategies. Throughan extensive empirical evaluation of 19 leading models, DSH-Bench un-covers previously obscured limitations in current approaches, establishingconcrete directions for future research and development.