研究大模型在人权议题上对不同族群的模糊回应,发现身份是影响模型态度的关键因素。
Hedging and Non-Affirmation: Quantifying LLM Alignment on Questions of Human Rights
- 构建框架量化模型对不同族群的人权表述模糊性
- 7个模型中4个对特定族群明显回避表态,身份影响远超其他因素
- 通过身份引导可有效减少模糊回应,适合关注模型偏见的开发者
回避和不确认是大语言模型在面对主观问题时表现出的倾向性行为,但在人权这类普遍适用的议题上却不可接受。本文提出系统性框架,评估六款商业大模型及一款开源模型在205个民族与无国家群体上的4738个非约束性提问中的表现。结果显示,7个模型中有4个在回应时表现出显著依赖于族群身份的回避与不确认行为。尽管冲突信号、主权状态(是否为无国家族群)或经济指标(如GDP)也影响模型行为,但其效应均弱于身份本身。该差异在重述问题后依然稳健。进一步使用开源模型测试,发现对族群身份进行引导(group steering)是最有效的去偏方法,且能抵抗下游遗忘问题。
原文摘要 · Abstract (English)
Hedging and non-affirmation are behaviors exhibited by large language models (LLMs) that limit the clear endorsement of specific statements. While these behaviors are desirable in subjective contexts, they are undesirable in the context of human rights - which apply unambiguously to all groups. We present a systematic framework to measure these behaviors in unconstrained LLM responses regarding various identity groups. We evaluate six large proprietary models as well as one open-weight LLM on 4738 prompts across 205 national and stateless ethnic identities and find that 4 out of 7 display hedging and non-affirmation that is significantly dependent on the identity of the group. While factors like conflict signals, sovereignty (whether identity is stateless), or economic indicators (GDP) also influence model behavior, their effect sizes are consistently weaker than the impact of identity itself. The systematic disparity is robust to methods of rephrasing the prompts. Since group identity is the strongest predictor of these behaviors, we use open-weight models to explore whether applying steering and orthogonalization techniques to these group identities can mitigate the rates of hedging and non-affirmation behaviors. We find that group steering is the most effective debiasing approach across query types and is robust to downstream forgetting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。