发现大模型隐性个性化行为的内部信号,可精准控制其影响。
Locating and Controlling Implicit Personalization in Large Language Models

- 通过对比有无暗示性线索的对话,定位到模型内部特定激活信号
- 该信号与推荐变化相关性高达r=0.87,多线索时信号叠加但输出不叠加
- 删除特定信号可有效抑制其影响,优于提示忽略,适合安全可控应用
大型语言模型(LLMs)在未明确用户身份的情况下,仍会根据隐含的人口统计学线索调整输出。尽管已有研究记录此现象,但其与模型内部激活之间的关联尚不清晰。我们通过五种LLM中匹配的带线索与中性对话,发现一个局部内部激活信号能追踪推荐变化,相关性最高达r=0.87。当多个线索共现时,其内部信号大致叠加,但输出变化并不简单相加。进一步发现,移除某一线索对应的内部信号可有效抑制其影响,通常比通过提示让模型忽略人口统计信息更有效,同时基本保持通用基准性能。然而,在保留共现维度影响的同时选择性消除某一维度的影响,仍高度依赖模型和属性类型。这些结果将隐性个性化行为与可分析、可因果操控的内部信号联系起来。
原文摘要 · Abstract (English)
Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r=0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension's influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。