LLM在模糊情境下会反映语言背后制度经验差异。
Do Large Language Models Encode Institutional Experience? Evidence from Cross-Linguistic Moral Reasoning Under Ambiguity
- 用跨语言道德困境测试模型是否继承制度性道德倾向
- 模糊情境中语言间道德分歧随真实制度质量差异增大
- 明确提示会抑制这种差异,适合研究伦理与语言关系者
大型语言模型(LLMs)在不同语言中的道德推理存在系统性差异,但其成因尚不明确。我们检验了语言可能编码其使用环境中的制度特征,使模型通过训练继承特定制度的道德先验。在涵盖九种语言、六种前沿大模型、两个预注册研究的实验中,我们考察了道德可接受性依赖于制度运作的情境。研究1显示,明确的制度背景设置未引发跨语言道德分歧,也未与语言社群间的制度差异相关。研究2引入制度关联但未明示的情境,在此条件下,跨语言道德分歧显著高于制度无关对照组,且除一个理论例外外,与真实世界制度质量差异正相关。明确提示再次削弱了这些效应。结果表明,制度经验可能以可检测的方式留在语言中,进而影响模型的道德推理,同时说明显性制度线索会抑制此类差异的表现。
原文摘要 · Abstract (English)
Large language models (LLMs) exhibit systematic differences in moral reasoning across languages, yet the source of this variation remains unclear. We test the hypothesis that languages encode aspects of the institutional environments in which they are spoken, allowing LLMs to inherit institution-specific moral priors through training. Across nine languages spanning a broad gradient of institutional quality, six frontier LLMs, and two preregistered studies, we examine moral dilemmas whose acceptability depends on institutional functioning. In Study 1, explicit institutional framing produced uniformly null results: cross-linguistic moral divergence did not increase in institutionally contingent scenarios, nor did it track institutional differences between language communities. In Study 2, we introduced institutionally ambiguous scenarios in which institutional stakes were present but not explicitly stated. Under these conditions, cross-linguistic moral divergence increased relative to institutionally inert controls and, with one theoretically informative exception, was associated with real-world institutional differences between language communities. Explicit framing again attenuated these effects. These findings suggest that institutional experience may leave detectable traces in language that shape LLM moral reasoning, while also indicating that explicit institutional cues can suppress the expression of those differences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。