Same Person, Different Depiction: Counterfactual Evaluation of Vision-Language Models on Individuals with Limb Deficiencies
Abstract
Vision-language models increasingly support everyday communication and decision-making from images. Their outputs shape how people talk about disability and can affect how people with disabilities are treated. Disability is common, but it is almost entirely absent from vision-side, controlled evaluations of model behavior. We ask a simple question: when a prompt is not about physical ability, should a model change its judgment just because visible limb-deficiency cues are present? A reliable model should base judgments and descriptions on task-relevant visual evidence, not on the mere presence of a disability cue. We introduce InclusiveCFImageBias, a benchmark of 924 images and more than 22k image-level queries. Each example contains two image versions that show the same person in the same scene, with and without visible limb-deficiency cues. We keep the background and activity fixed, and only edit cues such as a prosthesis or a residual limb. We query both versions with the same prompt and compare the outputs directly. Our evaluation spans both open-weight families (Qwen3-VL, DeepSeek-VL2, Gemma-3, and Ministral-3) and proprietary models (GPT-5 and Gemini-2.5). Many models shift both structured judgments and open-ended wording when limb-deficiency cues are visible. In our tests, these cues often push decisions toward more affirmative answers and higher ratings, and they make free responses less neutral and more evaluative. Higher scores are not automatically fairer. The key issue is that outputs change under a minimal, controlled visual substitution. InclusiveCFImageBias makes this cue-driven instability measurable. More importantly, it aims to draw community attention to disability-related VLM bias and provide a clear target for building more consistent disability-facing multimodal systems.