提出PACT框架,分析大模型在文化与个人偏好间的权衡行为
Whose Norms? Disentangling Cultural and Personal Alignment in Large Language Models

- 构建PACT框架,量化模型在文化规范与个人偏好间的决策倾向
- 模型行为受国家背景影响最大(7.8%),远超年龄(1%)和性别(0.7%)
- 适合关注社会判断、跨文化对齐及模型不确定性建模的研究者
大型语言模型在需平衡文化规范与个人偏好的社会决策场景中日益重要。现有研究多将文化对齐与个性化分开处理。本文提出PACT框架,用于评估模型在文化规范与个人偏好间的选择倾向。实验发现,不同模型对文化规范的遵守程度差异显著,国家背景的影响(7.8%)远大于年龄(1%)和性别(0.7%),且指令微调后变化不均一。五国人类研究显示,人类的文化遵循主要由情境国家决定,但在判断自身文化时一致性最低,体现文化内部多样性。人-模型对齐实验表明,模型可匹配多数选择,但无法捕捉响应分布与不确定性(最佳相关性仅0.24)。这些结果呼吁超越多数共识的对齐评估,以反映社会判断中的多元与分歧。
原文摘要 · Abstract (English)
Large language models are increasingly used for social decision-making situations that require balancing cultural norms with personal preferences. For example, a user preferring honesty might ask whether to correct a coworker publicly when local norms favor indirect feedback. Yet existing research studies cultural alignment and personalization largely separately. We introduce PACT, the Personal-Preference and Cultural-Norm Trade-off framework, which evaluates whether models choose to follow a cultural norm or allow personal preferences. We find that LLMs vary in how rigidly they enforce cultural norms, with behavior shifted more by country context (7.8%) than age (1%) and gender (0.7%) and shifting non-uniformly after instruction tuning. Furthermore, our five-country human study on PACT shows that culture-following in humans is mainly driven by scenario country, with the lowest agreement when participants judge their own cultural contexts, showing within-culture pluralism. Finally, human-LLM alignment experiments show that models can match majority choices, but fail to capture response distributions and uncertainty (with best correlations reaching only 0.24). Together, these findings motivate alignment evaluations that go beyond majority to capture cultural pluralism and disagreement in social judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。