arXiv:2506.16187cs.CLcs.AI2025-06AAAI被引 4

构建日语伦理理解评估数据集,揭示大模型在伦理推理上的短板。

JETHICS: Japanese Ethics Understanding Evaluation Dataset

  • 基于英文学术伦理框架构建7.8万条日语样本
  • GPT-4o平均得分仅0.7,日本本地模型约0.5
  • 适合评估AI伦理对齐能力的研究者使用

本文提出JETHICS,一个用于评估人工智能模型日语伦理理解能力的基准数据集。该数据集包含78,000个样本,沿用现有英文ETHICS数据集的构建方法,涵盖基于伦理与政治哲学规范理论的四个类别,以及代表常识道德的一个类别。在非专有大型语言模型(LLMs)和GPT-4o上的评估实验显示,即使是最先进的GPT-4o模型平均得分也仅为0.7左右,而表现最好的日本本地大模型得分约为0.5,表明当前大模型在日语伦理理解方面仍有巨大提升空间。

原文摘要 · Abstract (English)

In this work, we propose JETHICS, a Japanese dataset for evaluating ethics understanding of AI models. JETHICS contains 78K examples and is built by following the construction methods of the existing English ETHICS dataset. It includes four categories based normative theories and concepts from ethics and political philosophy; and one representing commonsense morality. Our evaluation experiments on non-proprietary large language models (LLMs) and on GPT-4o reveal that even GPT-4o achieves only an average score of about 0.7, while the best-performing Japanese LLM attains around 0.5, indicating a relatively large room for improvement in current LLMs.

伦理评估日语NLP大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。