拆解语言输入与推理语言对大模型道德判断的影响,发现推理语言作用更大。
Untangling Input Language from Reasoning Language: A Diagnostic Framework for Cross-Lingual Moral Alignment in LLMs
- 分离输入语言与推理语言,设计匹配与不匹配测试条件。
- 推理语言影响是输入语言的两倍,且近半模型存在上下文依赖性。
- 提出基于道德基础理论的解释框架,指导模型部署决策。
当大模型判断道德困境时,不同语言下结论是否不同?原因可能来自困境的语言本身,或模型推理所用语言。现有评估方法将两者混淆,仅测试语言一致的情况(如英文困境+英文推理)。本文提出新方法,分别操控两种语言,覆盖语言不匹配场景(如英文困境+中文推理),实现因素分解。为分析判断依据,引入道德基础理论进行解释,意外发现权威维度可细分为家庭相关与制度相关两类。在13个大模型上测试英汉道德判断,结果表明:(1)推理语言影响贡献的方差是输入语言的两倍;(2)该方法检测出近半模型的上下文依赖性,而标准评估未察觉;(3)构建诊断分类体系,转化为实际部署建议。代码与数据集已公开。
原文摘要 · Abstract (English)
When LLMs judge moral dilemmas, do they reach different conclusions in different languages, and if so, why? Two factors could drive such differences: the language of the dilemma itself, or the language in which the model reasons. Standard evaluation conflates these by testing only matched conditions (e.g., English dilemma with English reasoning). We introduce a methodology that separately manipulates each factor, covering also mismatched conditions (e.g., English dilemma with Chinese reasoning), enabling decomposition of their contributions. To study \emph{what} changes, we propose an approach to interpret the moral judgments in terms of Moral Foundations Theory. As a side result, we identify evidence for splitting the Authority dimension into a family-related and an institutional dimension. Applying this methodology to English-Chinese moral judgment with 13 LLMs, we demonstrate its diagnostic power: (1) the framework isolates reasoning-language effects as contributing twice the variance of input-language effects; (2) it detects context-dependency in nearly half of models that standard evaluation misses; and (3) a diagnostic taxonomy translates these patterns into deployment guidance. We release our code and datasets at https://anonymous.4open.science/r/CrossCulturalMoralJudgement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。