arXiv:2603.27338cs.AI2026-03

提出首个评估大模型道德判断编辑能力的基准数据集

CounterMoral: Editing Morals in Language Models

  • 构建跨伦理框架的道德判断编辑评测集
  • 验证现有编辑技术在道德内容修改上的有效性
  • 为打造符合人类价值观的语言模型提供评估工具

近年来,语言模型在事实信息编辑方面取得显著进展,但对道德判断这类关乎人类价值观的关键内容的修改仍缺乏关注。本文提出 CounterMoral,一个专门用于评估当前模型编辑技术在不同伦理框架下修改道德判断能力的基准数据集。我们对多个语言模型应用多种编辑方法,并进行系统评估。结果揭示了现有技术在道德内容调整上的局限性,为开发更符合人类伦理标准的语言模型提供了关键评估依据。

原文摘要 · Abstract (English)

Recent advancements in language model technology have significantly enhanced the ability to edit factual information. Yet, the modification of moral judgments, a crucial aspect of aligning models with human values, has garnered less attention. In this work, we introduce CounterMoral, a benchmark dataset crafted to assess how well current model editing techniques modify moral judgments across diverse ethical frameworks. We apply various editing techniques to multiple language models and evaluate their performance. Our findings contribute to the evaluation of language models designed to be ethical.

模型编辑道德对齐评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。