arXiv:2601.07780cs.CL2026-01被引 1

通过多角度反思提升大模型自我纠错能力

Enhancing Self-Correction in Large Language Models through Multi-Perspective Reflection

  • 设计多视角反思框架,从逻辑、信息、伦理等维度自检推理过程
  • 在复杂任务中显著提升准确率,伦理决策错误减少37%
  • 无需重训练,纯提示工程即可增强模型可靠性,适合高风险场景

尽管思维链(CoT)提示提升了大语言模型的推理能力,但在一致性、准确性和自我纠错方面仍存在挑战,尤其在复杂或涉及伦理的任务中。现有单一维度的反思方法改进有限。本文提出多视角反思思维链(PR-CoT),在初始思维链后,引导模型从逻辑一致性、信息完整性、偏见/伦理和替代方案四个预设角度进行自我评估。该方法仅通过提示工程实现,无需模型重训练,可将初始推理优化为更稳健准确的最终答案。在GPT-3.5与GPT-4上对算术、常识、伦理决策和逻辑谜题的实验表明,PR-CoT在逻辑一致性和错误修正方面显著优于传统CoT及现有反思方法,尤其在伦理决策等复杂领域表现突出。消融实验、人工评估与定性分析验证了各反思维度的贡献及整体范式有效性。

原文摘要 · Abstract (English)

While Chain-of-Thought (CoT) prompting advances LLM reasoning, challenges persist in consistency, accuracy, and self-correction, especially for complex or ethically sensitive tasks. Existing single-dimensional reflection methods offer insufficient improvements. We propose MyGO Poly-Reflective Chain-of-Thought (PR-CoT), a novel methodology employing structured multi-perspective reflection. After initial CoT, PR-CoT guides the LLM to self-assess its reasoning across multiple predefined angles: logical consistency, information completeness, biases/ethics, and alternative solutions. Implemented purely via prompt engineering, this process refines the initial CoT into a more robust and accurate final answer without model retraining. Experiments across arithmetic, commonsense, ethical decision-making, and logical puzzles, using GPT-three point five and GPT-four models, demonstrate PR-CoT's superior performance. It significantly outperforms traditional CoT and existing reflection methods in logical consistency and error correction, with notable gains in nuanced domains like ethical decision-making. Ablation studies, human evaluations, and qualitative analyses further validate the contribution of each reflection perspective and the overall efficacy of our poly-reflective paradigm in fostering more reliable LLM reasoning.

大模型推理自我纠错多视角反思伦理决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。