不同身份提示导致大模型响应不一致,暴露评估漏洞
Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias
- 用多种身份线索测试模型响应差异
- 同一群体线索引发的响应变化仅部分重叠
- 建议多线索评估以避免误判偏差
基于身份线索的评估广泛用于研究大语言模型(LLMs)如何根据显式身份属性调整回应。该方法通常依赖单一线索(如姓名)作为群体归属的代理,隐含假设不同线索可互换使用。我们在涵盖1480万次提示的真实咨询交互中,考察了美国语境下种族与性别因素的影响。结果发现,同一群体的不同线索引发的模型响应变化仅部分重叠,导致个性化结论不一致;偏差结论也极不稳定,组间差异的大小和方向随线索变化而波动。进一步分析表明,这些不一致性源于线索-群体关联强度差异以及线索内嵌的语言特征对模型行为的影响。总体而言,我们的研究揭示:大模型的身份响应并非线索无关的类别级参数,而是高度依赖于身份提示方式,反映的是对语言信号的响应,而非稳定的身份类别。因此,我们呼吁采用多线索、机制感知的评估范式,以建立关于大模型身份差异的稳健且可解释的结论。
原文摘要 · Abstract (English)
Demographic cue-based evaluation is widely used to study how large language models (LLMs) adapt their responses to signaled demographic attributes within and across groups. This approach typically relies on a single cue (e.g., names) as a proxy for group membership, implicitly treating different cues as interchangeable operationalizations of a single underlying identity-conditioned behavior. We test this assumption in realistic advice-seeking interactions spanning 14.8 million prompts, focusing on race and gender in a U.S. context. We find that cues for the same group induce only partially overlapping changes in model responses, yielding inconsistent conclusions about personalization, while bias conclusions are unstable, with both magnitude and direction of group differences varying across cues. We further show that these inconsistencies reflect differences in cue-group association strength and linguistic features bundled within cues that shape model responses. Together, our findings suggest that demographic conditioning in LLMs is not a cue-invariant category-level parameter but depends fundamentally on how identity is cued, reflecting responses to linguistic signals rather than stable demographic categories. We therefore call for multi-cue, mechanism-aware evaluations as a foundation for robust and interpretable claims about demographic variation in LLM responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。