arXiv:2609.05993cs.CL2026-09

大模型为适应文化偏好,会抹去个体差异,用刻板印象提升准确率。

Alignment by Stereotyping: How LLMs Sacrifice Individual Distinctiveness for Cultural Adaptation

  • 通过用户人口统计信息增强文化适应性,但导致个体特征被群体均值替代。
  • 7个模型在世界价值观调查中显示,性能越强越倾向压缩个体差异,远超人类水平。
  • 对话中分散人口统计信号可部分抑制刻板印象,适合关注个性化交互的研究者。

大型语言模型在个性化交互中日益普及,通过用户档案进行人口统计条件化是常见的文化适应策略。我们探讨这一方法是否真正服务个体,还是通过消除个体差异来提高准确性。基于世界价值观调查数据,研究七种模型(包括前沿的GPT-5.1)发现,人口统计信息虽提升了多数模型的价值对齐准确率,但系统性地牺牲了个体独特性。模型将回答向人口群体中心点靠拢,这种行为称为‘刻板印象式对齐’。置换检验(10,000次置换,6个人口属性,7个模型)证实,表现最优的模型对个体的压缩程度远超人类基线;家庭内部尺度扩展放大了这一权衡,同时削弱了内在文化理解能力。使用经真实人机对话验证的合成对话数据集(来自PRISM,Kirk et al., 2024),进一步表明在对话轮次中分散人口统计信号,相比集中标签能部分抑制原型检索,该发现已在真实对话中通过PRISM验证,但需更大规模复现。

原文摘要 · Abstract (English)

Large language models are increasingly deployed for personalized interaction, and demographic conditioning via user profiles is a widely adopted strategy for cultural adaptation. We ask whether this approach genuinely serves individual users or achieves accuracy by erasing individual distinctiveness. Studying seven models including frontier GPT-5.1 on the World Values Survey, we find that demographic profiles improve value alignment accuracy for most models, but at a systematic cost to individuality. That is, models pull responses toward demographic group centroids rather than preserving individual differences, a behavioral pattern we term alignment by stereotyping. Permutation tests (10,000 permutations, six demographic attributes, seven models) certify that top-performing models compress individuals far above the human baseline; within-family scaling amplifies this tradeoff while degrading intrinsic cultural understanding. Using a synthetic dialogue dataset validated on real human-chatbot conversations from PRISM (Kirk et al., 2024), we further show that distributing demographic signals across conversational turns partially suppresses prototype retrieval compared to compact demographic labels, a finding validated on real conversations via PRISM but requiring replication at larger scale.

大模型对齐刻板印象个性化交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。