arXiv:2410.20513cs.CL2024-10被引 1

大模型不具备内在的自我纠错能力,尤其在道德判断上。

Self-correction is Not An Innate Capability in Language Models

  • 通过自辨任务分析模型道德敏感性,发现其缺乏内在道德判断力。
  • 机制分析表明,思维链与外部反馈难以有效结合,导致纠错失效。
  • 适合关注大模型伦理安全与可信推理的研究者阅读。

尽管大语言模型(LLMs)的自我纠错能力受到广泛关注,但对其有效性尚无定论。以往研究多聚焦于内在自我纠错,而对内部知识与外部反馈之间的相互作用等外在自我纠错机制则关注不足。本文旨在深入探究道德自我纠错的底层机制,回答一个根本问题:道德自我纠错是否为大模型的固有能力?具体开展两项研究:(1) 基于自辨任务的行为分析,考察大模型的道德敏感性;(2) 通过对隐藏状态的机制分析,探究思维链(Chain-of-Thought, CoT)与外部反馈等关键组件如何协同促进道德自我纠错。基于行为与机制两方面的实证证据,我们证明:道德自我纠错并非大模型的固有能力,它们既缺乏道德敏感性,也无法在自我纠错过程中有效整合外部反馈。

原文摘要 · Abstract (English)

Although there has been growing interest in the self-correction capability of Large Language Models (LLMs), there are varying conclusions about its effectiveness. Prior research has largely concentrated on intrinsic self-correction, extrinsic self-correction, particularly the interplay between internal knowledge and external feedback, remains underexplored. In this paper, we aim to comprehensively investigate the underlying mechanism of moral self-correction by addressing a fundamental question: is moral self-correction an innate capability of LLMs? Specifically, we conduct: (1) a behavioral analysis of LLMs' moral sensitivity based on a self-distinguishing task; and (2) a mechanistic analysis of the hidden states to examine how key components of self-correction, such as Chain-of-Thought (CoT) and external feedback, interact to facilitate moral self-correction. Drawing on empirical evidence from both behavioral and mechanistic analyses, we demonstrate that moral self-correction is not an inherent capability of LLMs, as they are neither morally sensitive nor able to effectively incorporate external feedback during the self-correction process.

大模型自我纠错道德判断机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。