arXiv:2510.07290cs.CLcs.LG2025-10被引 1

大模型能通过多轮自我修正实现道德回应的稳定优化。

On the Convergence of Moral Self-Correction in Large Language Models

  • 通过多轮指令激活模型内部道德概念,逐步降低不确定性。
  • 连续修正后模型输出趋于稳定,性能收敛至最优水平。
  • 适合研究大模型伦理对齐与自我改进机制的学者参考。

大型语言模型(LLMs)在收到指令时可自我修正其回应,这一能力称为自校正。当指令仅提供抽象目标而无具体问题描述时,模型需依赖内部知识进行改进,称为内在自校正。尽管内在自校正已在多种应用中表现成功,其作用机制仍不明确。本文聚焦于大模型的道德自校正,揭示其关键特征:通过多轮交互实现性能收敛;并对其收敛行为进行机制分析。实验结果表明,持续注入的自校正指令会激活模型中的道德概念,减少模型不确定性,使被激活的道德概念在多轮迭代中趋于稳定,从而导致性能收敛。本研究展示了道德自校正的强大潜力,证明其具备理想的表现收敛特性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are able to improve their responses when instructed to do so, a capability known as self-correction. When instructions provide only a general and abstract goal without specific details about potential issues in the response, LLMs must rely on their internal knowledge to improve response quality, a process referred to as intrinsic self-correction. The empirical success of intrinsic self-correction is evident in various applications, but how and why it is effective remains unknown. Focusing on moral self-correction in LLMs, we reveal a key characteristic of intrinsic self-correction: performance convergence through multi-round interactions; and provide a mechanistic analysis of this convergence behavior. Based on our experimental results and analysis, we uncover the underlying mechanism of convergence: consistently injected self-correction instructions activate moral concepts that reduce model uncertainty, leading to converged performance as the activated moral concepts stabilize over successive rounds. This paper demonstrates the strong potential of moral self-correction by showing that it exhibits a desirable property of converged performance.

大模型自我修正道德对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。