不同提示框架显著影响大模型文化价值观输出,第三视角预测最有效。
Personalization, Personas, and Forecasting in Value Alignment

- 用13个语言-国家组合测试四种模型,对比用户、角色、第三人称提示差异。
- 21,008条响应显示,第三人称提示在三个模型上对齐人类分布最优。
- 宗教、性别角色等核心价值维度更易对齐,制度信任类问题仍难改善。
大模型行为可能通过多种方式受人类身份影响:适配用户、扮演群体角色或预测人们对价值问题的回答。我们基于世界价值观调查(WVS)测试这些表述是否可互换。在13个语言-国家组合中,评估GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash和Qwen3-235B在101个源自WVS的问题上的表现,对比仅语言基线与用户-国家、人格-国家、第三人称提示的差异。在21,008条模型响应中,提示框架是文化对齐的一阶决定因素:国家线索常显著改变答案,但并非所有变化都趋向人类响应分布。第三人称预测在四个模型中的三个实现最强方向性对齐,而个性化和角色扮演效果较弱或不稳定。对齐增益集中于显著的价值维度,如宗教性、性别角色和以工作为导向的物质价值观,而制度信任和民主相关问题仍具挑战。结果表明,提示框架在文化价值探测中并非装饰性选择,它会实质性改变模型行为和测得的对齐程度。
原文摘要 · Abstract (English)
LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions. We test whether these framings are interchangeable using the World Values Survey (WVS). We evaluate GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Qwen3-235B on 101 WVS-derived questions across 13 language-country slices, comparing a language-only baseline with user-country, persona-country, and third-person prompts. Across 21,008 model-response rows, prompt framing is a first-order determinant of cultural alignment: country cues often shift answers substantially, but not all shifts move toward matched human response distributions. Third-person forecasting yields the strongest directional alignment for three of the four hosted models, while personalization and role-play are weaker or less stable. Alignment gains concentrate on salient value dimensions such as religiosity, gender roles, and work-oriented material values, whereas institutional trust and democracy-related questions remain difficult. These results show that prompt framing is not a cosmetic choice in cultural value elicitation; it changes both model behavior and measured alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。