arXiv:2508.12754cs.AI2025-08被引 3

评估大模型能否像人类道德助手一样进行深层推理,而不仅是判断对错。

Beyond Ethical Alignment: Evaluating LLMs as Artificial Moral Assistants

  • 提出道德助手应具备演绎与归纳推理能力,超越单纯价值对齐。
  • 测试发现主流大模型在归纳推理上表现薄弱,存在明显短板。
  • 适合关注AI伦理、道德推理与模型可解释性的研究者阅读。

大型语言模型(LLMs)的兴起引发了对其道德能力的广泛关注。尽管已有大量工作致力于将模型与人类道德价值观对齐,但现有评估基准仍过于表面,主要依赖最终的道德判断,而非显式的道德推理过程。本文旨在推进对LLMs道德能力的研究,考察其作为人工道德助手(Artificial Moral Assistants, AMAs)的潜力——这类系统在哲学文献中被设想为支持人类道德思辨的工具。我们提出,成为合格的AMA不仅需识别道德困境,更需主动推理,处理对齐阶段未包含的价值冲突。基于哲学理论,我们构建了一个新的行为框架,界定关键特征如演绎与归纳道德推理,并据此设计评估基准,测试多个主流开源大模型。结果表明模型间表现差异显著,尤其在归纳推理方面存在持续缺陷。本研究连接了哲学理论与实际AI评估,强调需专门策略提升模型的道德推理能力。代码已开源。

原文摘要 · Abstract (English)

The recent rise in popularity of large language models (LLMs) has prompted considerable concerns about their moral capabilities. Although considerable effort has been dedicated to aligning LLMs with human moral values, existing benchmarks and evaluations remain largely superficial, typically measuring alignment based on final ethical verdicts rather than explicit moral reasoning. In response, this paper aims to advance the investigation of LLMs' moral capabilities by examining their capacity to function as Artificial Moral Assistants (AMAs), systems envisioned in the philosophical literature to support human moral deliberation. We assert that qualifying as an AMA requires more than what state-of-the-art alignment techniques aim to achieve: not only must AMAs be able to discern ethically problematic situations, they should also be able to actively reason about them, navigating between conflicting values outside of those embedded in the alignment phase. Building on existing philosophical literature, we begin by designing a new formal framework of the specific kind of behaviour an AMA should exhibit, individuating key qualities such as deductive and abductive moral reasoning. Drawing on this theoretical framework, we develop a benchmark to test these qualities and evaluate popular open LLMs against it. Our results reveal considerable variability across models and highlight persistent shortcomings, particularly regarding abductive moral reasoning. Our work connects theoretical philosophy with practical AI evaluation while also emphasising the need for dedicated strategies to explicitly enhance moral reasoning capabilities in LLMs. Code available at https://github.com/alessioGalatolo/AMAeval

道德推理大模型评估伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。