arXiv:2508.15831cs.CLcs.AI2025-08ICCV被引 8

检测大模型如何因残疾相关表述产生偏见,发现模型常无根据猜测用户身份。

Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs

  • 用9类残疾+6个业务场景构建平衡测试集,评估8个大模型的推断行为
  • 97%情况下模型会无依据猜测用户性别、教育等5项属性,且受残疾描述显著影响
  • 模型越大越敏感于残疾线索,建议引入拒绝回答和反事实训练缓解偏见

大型语言模型常仅凭用户措辞就推断其人口属性,可能导致偏见响应,即使未提供明确信息。残疾相关线索在其中的作用仍不清晰。本文首次系统审计了8个参数量从30亿到720亿的先进指令微调模型在残疾情境下的偏见表现。通过一个平衡模板语料库,将9类残疾类别与6个真实业务领域配对,让每个模型在中性与残疾感知条件下预测5个人口属性(性别、社会经济地位、教育程度、文化背景、所在地)。结果显示,模型在高达97%的提示中会给出明确的人口推断,表现出强随意性且缺乏合理依据;残疾上下文显著改变预测分布,行业背景可进一步放大偏差。更大的模型反而更敏感于残疾线索,也更易出现偏见推理,表明规模并非缓解刻板印象的保障。研究揭示了能力主义与其他歧视之间的深层交集,指出当前对齐策略的关键盲区。我们发布评估框架与结果,倡导包容残疾的基准测试,并建议引入拒答校准与反事实微调以抑制不当推断。代码与数据将在论文接受后公开。

原文摘要 · Abstract (English)

Large Language Models (LLMs) routinely infer users demographic traits from phrasing alone, which can result in biased responses, even when no explicit demographic information is provided. The role of disability cues in shaping these inferences remains largely uncharted. Thus, we present the first systematic audit of disability-conditioned demographic bias across eight state-of-the-art instruction-tuned LLMs ranging from 3B to 72B parameters. Using a balanced template corpus that pairs nine disability categories with six real-world business domains, we prompt each model to predict five demographic attributes - gender, socioeconomic status, education, cultural background, and locality - under both neutral and disability-aware conditions. Across a varied set of prompts, models deliver a definitive demographic guess in up to 97\% of cases, exposing a strong tendency to make arbitrary inferences with no clear justification. Disability context heavily shifts predicted attribute distributions, and domain context can further amplify these deviations. We observe that larger models are simultaneously more sensitive to disability cues and more prone to biased reasoning, indicating that scale alone does not mitigate stereotype amplification. Our findings reveal persistent intersections between ableism and other demographic stereotypes, pinpointing critical blind spots in current alignment strategies. We release our evaluation framework and results to encourage disability-inclusive benchmarking and recommend integrating abstention calibration and counterfactual fine-tuning to curb unwarranted demographic inference. Code and data will be released on acceptance.

模型偏见残疾歧视大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。