arXiv:2502.14083cs.CL2025-02ACL被引 18

构建多语言道德推理数据集,揭示大模型在跨文化道德判断中的能力与局限。

Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral

  • 设计统一数据集UniMoral,整合六种语言的道德困境与多维度标注。
  • 三款大模型在四项任务中表现不一,显示隐含道德语境提升推理能力。
  • 适用于研究跨文化道德认知、模型伦理评估及多语言NLP的学者。

道德推理是受个体经历和文化背景影响的复杂认知过程,对计算分析构成独特挑战。尽管自然语言处理(NLP)提供了研究该现象的潜力,现有研究缺乏一致性,使用分散的数据集和任务,仅考察道德推理的孤立方面。我们通过UniMoral弥合这一空白,该数据集整合了心理学基础与社交媒体来源的道德困境,涵盖行动选择、伦理原则、影响因素和后果等标注,并附有标注者道德与文化背景信息。考虑到道德推理的文化相对性,UniMoral覆盖阿拉伯语、中文、英语、印地语、俄语和西班牙语六种语言,捕捉多样社会文化情境。我们通过基准评估三种大语言模型(LLMs)在四项任务上的表现:行动预测、道德类型分类、因素归因分析和后果生成。关键发现表明,虽然隐含的道德语境能增强大模型的道德推理能力,但仍需更专门的方法以进一步推进其性能。

原文摘要 · Abstract (English)

Moral reasoning is a complex cognitive process shaped by individual experiences and cultural contexts and presents unique challenges for computational analysis. While natural language processing (NLP) offers promising tools for studying this phenomenon, current research lacks cohesion, employing discordant datasets and tasks that examine isolated aspects of moral reasoning. We bridge this gap with UniMoral, a unified dataset integrating psychologically grounded and social-media-derived moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, alongside annotators' moral and cultural profiles. Recognizing the cultural relativity of moral reasoning, UniMoral spans six languages, Arabic, Chinese, English, Hindi, Russian, and Spanish, capturing diverse socio-cultural contexts. We demonstrate UniMoral's utility through a benchmark evaluations of three large language models (LLMs) across four tasks: action prediction, moral typology classification, factor attribution analysis, and consequence generation. Key findings reveal that while implicitly embedded moral contexts enhance the moral reasoning capability of LLMs, there remains a critical need for increasingly specialized approaches to further advance moral reasoning in these models.

道德推理多语言大模型评估数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。