模型看似公平,实则隐性偏见严重,单一测试无法发现真相。
Redirected, Not Removed: Task-Dependent Stereotyping Reveals the Limits of LLM Alignments
- 设计9类偏见分类与7项任务,覆盖显性与隐性测试
- 同一模型在不同任务中偏见分差达0.43,显示任务依赖性
- 对边缘群体避责但为优势群体贴正向标签,安全对齐不对称
语言模型的偏见程度取决于提问方式。一个拒绝在种姓间选领导的模型,在填空任务中却会将上层种姓与纯洁关联,下层种姓与卫生差关联。单任务基准忽略这一点,仅捕捉偏见的片面表现。我们提出涵盖9类偏见(包括种姓、语言、地理等未被充分研究维度)的层级分类体系,并构建7项评估任务,从显性决策到隐性联想全覆盖。通过约4.5万条提示语审计7个商用及开源大模型,发现三类系统性模式:第一,偏见具有任务依赖性——模型在显性任务中抵制刻板印象,但在隐性任务中重现,同一模型对同一群体的刻板印象得分在不同任务间差异高达0.43;第二,安全对齐呈非对称性——模型拒绝给边缘群体贴负面标签,却自由将正面特质赋予优势群体;第三,未被充分研究的偏见维度表现出最强刻板印象,表明对齐努力更追随评测覆盖范围而非危害严重性。结果表明,单基准审计会系统性误判模型偏见,当前对齐实践掩盖而非消解表征伤害。
原文摘要 · Abstract (English)
How biased is a language model? The answer depends on how you ask. A model that refuses to choose between castes for a leadership role will, in a fill-in-the-blank task, reliably associate upper castes with purity and lower castes with lack of hygiene. Single-task benchmarks miss this because they capture only one slice of a model's bias profile. We introduce a hierarchical taxonomy covering 9 bias types, including under-studied axes like caste, linguistic, and geographic bias, operationalized through 7 evaluation tasks that span explicit decision-making to implicit association. Auditing 7 commercial and open-weight LLMs with \textasciitilde45K prompts, we find three systematic patterns. First, bias is task-dependent: models counter stereotypes on explicit probes but reproduce them on implicit ones, with Stereotype Score divergences up to 0.43 between task types for the same model and identity groups. Second, safety alignment is asymmetric: models refuse to assign negative traits to marginalized groups, but freely associate positive traits with privileged ones. Third, under-studied bias axes show the strongest stereotyping across all models, suggesting alignment effort tracks benchmark coverage rather than harm severity. These results demonstrate that single-benchmark audits systematically mischaracterize LLM bias and that current alignment practices mask representational harm rather than mitigating it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。