LLM会通过刻板印象推断用户身份,导致少数群体回应质量下降。
Reading Between the Prompts: How Stereotypes Shape LLM's Implicit Personalization
- 用合成对话测试模型如何从隐含线索推断用户身份。
- 模型在用户明确声明身份后仍受刻板印象影响,偏差持续存在。
- 通过线性探测干预内部表征,可有效减轻刻板印象偏差。
生成式大语言模型(LLMs)会从对话中的细微线索推断用户人口属性,这一现象称为隐式个性化。已有研究显示,即使未提供显式人口信息,此类推断也可能导致少数群体获得较低质量的回复。本文通过受控的合成对话,结合模型内部表征分析与针对性问题的回答,系统探究了LLMs对刻板印象线索的响应机制。结果表明,模型确实会基于这些刻板信号推断人口属性,且对部分群体而言,这种推断即使在用户明确声明不同身份后仍持续存在。最后,我们证明可通过训练线性探测器干预模型内部表示,将表征引导至用户明确陈述的身份,从而有效缓解此类由刻板印象驱动的隐式个性化。研究强调了提升大模型用户身份表征透明度与可控性的必要性。
原文摘要 · Abstract (English)
Generative Large Language Models (LLMs) infer user's demographic information from subtle cues in the conversation -- a phenomenon called implicit personalization. Prior work has shown that such inferences can lead to lower quality responses for users assumed to be from minority groups, even when no demographic information is explicitly provided. In this work, we systematically explore how LLMs respond to stereotypical cues using controlled synthetic conversations, by analyzing the models' latent user representations through both model internals and generated answers to targeted user questions. Our findings reveal that LLMs do infer demographic attributes based on these stereotypical signals, which for a number of groups even persists when the user explicitly identifies with a different demographic group. Finally, we show that this form of stereotype-driven implicit personalization can be effectively mitigated by intervening on the model's internal representations using a trained linear probe to steer them toward the explicitly stated identity. Our results highlight the need for greater transparency and control in how LLMs represent user identity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。