BBQ-V: Benchmarking Visual Stereotype Bias in Large Multimodal Models
Abstract
Stereotype biases in Large Multimodal Models (LMMs) per-petuate harmful societal prejudices, undermining the fairness and equityof AI applications. As LMMs grow increasingly influential, addressingand mitigating inherent biases related to stereotypes, harmful generations,and ambiguous assumptions in real-world scenarios has become essential.However, existing datasets evaluating stereotype biases in LMMs oftenlack diversity, rely on synthetic images, and often have single-actor im-ages, leaving a gap in bias evaluation for real-world visual contexts. Toaddress the gap in bias evaluation using real images, we introduce theBBQ-Vision (BBQ-V), the most comprehensive framework for assessingstereotype biases across nine diverse categories and 50 sub-categories withreal and multi-actor images. BBQ-V benchmark contains 14,144 image-question pairs and rigorously evaluates LMMs through carefully curated,visually grounded scenarios, challenging them to reason accurately aboutvisual stereotypes. It offers a robust evaluation framework featuring real-world visual samples, image variations, and open-ended question formats.BBQ-V enables a precise and nuanced assessment of a model’s reasoningcapabilities across varying levels of difficulty. Through rigorous testingof 19 state-of-the-art open-source (general-purpose and reasoning) andclosed-source LMMs, we highlight that these top-performing models areoften biased on several social stereotypes, and demonstrate that the think-ing models induce more bias in the reasoning chains. This benchmarkrepresents a significant step toward fostering fairness in AI systems andreducing harmful biases, laying the groundwork for equitable and sociallyresponsible LMMs. Dataset and evaluation code are available here.