构建跨文化多语言道德推理框架,让AI更懂不同文化的道德判断。
MET: Theory-Grounded and Culture-Aware Multilingual Moral Reasoning

- 基于心理学与哲学理论设计双阶段提示,支持跨文化情境下的本地化推理。
- 在多语言数据集上平均提升3.71分,最大增益达12.94分,显著改善道德决策能力。
- 无需外部标注即可自蒸馏训练,适合多语言伦理应用与低资源场景研究者。
语言模型越来越多地用于跨语言与文化背景的道德决策,但现有研究在三个方面存在忽视:1)多语言评估基准采用直接翻译,未适配文化特定内容;2)推理方法依赖静态英语中心的模板,缺乏道德理论支撑;3)道德决策训练通常需昂贵的强模型监督或人工标注。本文提出三项贡献:首先,构建MCLASH多语言道德决策基准,捕捉不同语言中的文化情境化道德直觉与社会规范;其次,提出MET(Multilingual Ethics with Theory-grounded reasoning),一种基于专家标注理论基础的两步提示方法:模型先选择情境与文化相关的理由,再以用户母语进行推理;第三,引入MET-D(MET-Distillation),通过无外部监督的自蒸馏训练增强第二步推理能力。实验显示,MET-D在三种不同规模与架构的模型(Qwen3-4B、Qwen3-8B、Gemma3-4B)上均提升性能,于MCLASH平均提升3.71分,于MMoralExceptQA提升4.23分,其中马来语上最高达12.94分。此外,母语推理能力平均提升62.13分,且有益理由呈现系统性文化差异。这些成果为文化对齐、理论驱动的多语言道德推理开辟新路径。
原文摘要 · Abstract (English)
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。