arXiv:2507.00985cs.CL2025-07EMNLP被引 7

发现大模型道德自纠中的隐性思维捷径,提出改进方案。

Discourse Heuristics For Paradoxically Moral Self-Correction

  • 分析语料发现道德自纠依赖隐性思维捷径
  • 捷径导致自纠与自诊能力难以同时提升
  • 基于精选数据集优化,提升道德对齐效果

道德自纠已成为对齐大语言模型输出与人类道德价值观的有前景方法。然而,该技术面临两大悖论:其一,尽管实证与理论证据支持其有效性,但自纠能力仅停留在表层;其二,尽管模型可诊断输出中的不道德内容,却难以在自纠过程中识别道德不一致的根本原因。为深入理解并解决这些悖论,我们分析了用于增强道德自纠的微调语料中的话语构建,揭示出有效构建背后的启发式规则。研究证明,道德自纠依赖于反映启发式捷径的话语结构,而这些捷径的存在会导致在同时提升自纠与自诊能力时产生不一致性。基于此,我们提出一种通过利用精选数据集的启发式来改进道德自纠的方法,并强调该能力在情境化上下文和模型规模上的泛化挑战。

原文摘要 · Abstract (English)

Moral self-correction has emerged as a promising approach for aligning the output of Large Language Models (LLMs) with human moral values. However, moral self-correction techniques are subject to two primary paradoxes. First, despite empirical and theoretical evidence to support the effectiveness of self-correction, this LLM capability only operates at a superficial level. Second, while LLMs possess the capability of self-diagnosing immoral aspects of their output, they struggle to identify the cause of this moral inconsistency during their self-correction process. To better understand and address these paradoxes, we analyze the discourse constructions in fine-tuning corpora designed to enhance moral self-correction, uncovering the existence of the heuristics underlying effective constructions. We demonstrate that moral self-correction relies on discourse constructions that reflect heuristic shortcuts, and that the presence of these heuristic shortcuts during self-correction leads to inconsistency when attempting to enhance both self-correction and self-diagnosis capabilities jointly. Based on our findings, we propose a solution to improve moral self-correction by leveraging the heuristics of curated datasets. We also highlight the generalization challenges of this capability, particularly in terms of learning from situated context and model scales.

大模型对齐道德推理语料分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。