提出新方法评估带安全防护的视觉语言模型偏见,发现所有模型都会受用户性别影响输出。
Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models

- 不直接问人物属性,改用虚构故事等任务隐含传递用户性别信息
- 在20个主流模型上测试,发现男女用户下角色描述明显不同
- 适合关注模型社会偏见与安全机制的开发者和研究者
我们提出一种面向强安全防护环境下大型视觉语言模型(LVLMs)的社会偏见评估方法。现有基准依赖要求模型推断图像中人物属性的提示(如“此人是总裁还是秘书?”),但具备强防护机制的模型(如GPT、Claude)常拒绝此类请求,导致评估不可靠。为此,我们改变评估范式:将任务与人物解耦,采用不涉及人物属性的提示(如“写一个关于虚构人物的虚构故事”),并将图像作为临时用户信息隐含传递人口统计线索,比较不同用户群体下的输出差异。该方法在故事生成、术语解释和考试式问答三个任务中均避免拒绝行为,实现可靠偏见测量。对20个近期开放源代码与专有模型的应用结果显示,所有模型均在非相关任务中不当使用用户人口统计信息——例如,男性用户对应的虚构角色多为机械师,女性用户则多为护士。尽管仍存在偏见,专有模型如GPT-5的偏差低于开源模型。我们分析了这一差距的潜在原因,讨论持续模型监控与优化可能是降低偏见的关键因素。
原文摘要 · Abstract (English)
We propose a societal bias evaluation method for large vision-language models (LVLMs) in the era of strong safety guardrails. Existing benchmarks rely on prompts that ask models to infer attributes of people in images (e.g., "Is this person a CEO or a secretary?"). However, we find that LVLMs with strong guardrails, such as GPT and Claude, often refuse these prompts, making evaluations unreliable. To address this, we change the prior evaluation paradigm by decoupling the task from the depicted person: instead of inferring person's attributes, we use prompts that do not ask about the person (e.g., "Write a fictional story about an imaginary person.") and attach the image as provisional user information to implicitly provide demographic cues, then compare outputs across user demographics. Instantiated across three tasks --- story generation, term explanation, and exam-style QA --- our method avoids refusals even in guardrailed LVLMs, enabling reliable bias measurement. Applying it to 20 recent LVLMs, both open-source and proprietary, we find that all models undesirably use user demographic information in person-irrelevant tasks; for instance, characters in stories are often portrayed as mechanic for male users and nurse for female users. Although still biased, proprietary models like GPT-5 show lower bias than open-source ones. We analyze potential factors behind this gap, discussing continuous model monitoring and improvement as a possible contributor for reducing bias.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。