arXiv:2506.19073cs.CL2025-06EMNLP被引 8

构建多语言仇恨言论道德推理数据集,揭示大模型在道德判断上的严重不足

MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation

  • 基于道德基础理论构建多跳推理的跨语言评估数据集
  • 大模型仇恨检测表现尚可(F1最高0.836),但道德判断准确率低于0.35
  • 低资源语言的解释对齐差,凸显文化差异下的伦理理解短板

确保大型语言模型(LLMs)具备道德推理能力日益重要,尤其在涉及社会敏感任务时。然而现有评估基准存在两大缺陷:缺乏支持道德分类的标注,限制了透明性与可解释性;且以英语为主,难以评估跨文化语境下的道德推理。本文提出MFTCXplain,一个基于道德基础理论、通过多跳仇恨言论解释来评估LLM道德推理能力的多语言基准数据集。该数据集包含4种语言(葡萄牙语、意大利语、波斯语、英语)共3,000条推文,每条均标注二元仇恨言论标签、道德类别及文本片段级理由。结果表明,大模型输出与人工标注在道德推理上存在显著偏差。尽管其在仇恨言论检测中表现良好(最高F1达0.836),但在预测道德情感方面表现极弱(F1<0.35)。此外,理由对齐在低资源语言中仍不理想。

原文摘要 · Abstract (English)

Ensuring the moral reasoning capabilities of Large Language Models (LLMs) is a growing concern as these systems are used in socially sensitive tasks. Nevertheless, current evaluation benchmarks present two major shortcomings: a lack of annotations that justify moral classifications, which limits transparency and interpretability; and a predominant focus on English, which constrains the assessment of moral reasoning across diverse cultural settings. In this paper, we introduce MFTCXplain, a multilingual benchmark dataset for evaluating the moral reasoning of LLMs via multi-hop hate speech explanation using the Moral Foundations Theory. MFTCXplain comprises 3,000 tweets across Portuguese, Italian, Persian, and English, annotated with binary hate speech labels, moral categories, and text span-level rationales. Our results show a misalignment between LLM outputs and human annotations in moral reasoning tasks. While LLMs perform well in hate speech detection (F1 up to 0.836), their ability to predict moral sentiments is notably weak (F1 < 0.35). Furthermore, rationale alignment remains limited mainly in underrepresented languages. Our findings show the limited capacity of current LLMs to internalize and reflect human moral reasoning

道德推理多语言仇恨言论可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。