arXiv:2509.07676cs.AI2025-09

通过用户反馈触发纠错,实现更深度的多路径推理。

Unleashing the True Potential of LLMs: A Feedback-Triggered Self-Correction with Long-Term Multipath Decoding

  • 仅在收到负面反馈时才触发重生成,避免错误自我评估
  • 采用延迟评估的多路径解码,探索多种推理路径
  • 在数学与代码生成任务中显著优于现有自纠错方法

大型语言模型在各类任务中表现优异,但推理过程中生成错误内容的问题仍难以解决。现有自纠错方法受限于缺乏可靠的错误定位信号,以及传统逐词预测带来的推理深度不足。为此,我们提出反馈触发重生成(FTR)框架,结合用户反馈与增强的解码机制:仅当收到负面反馈时才启动响应重生成,避免错误自我评估导致的错误传播,同时保留原始正确输出。此外,引入长期多路径(LTM)解码,通过延迟序列评估系统探索多条推理轨迹,有效克服标准逐词预测的短视决策问题。在数学推理与代码生成基准上的大量实验表明,该框架在性能上持续且显著优于当前最优的提示式自纠错方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have achieved remarkable performance across diverse tasks, yet their susceptibility to generating incorrect content during inference remains a critical unsolved challenge. While self-correction methods offer potential solutions, their effectiveness is hindered by two inherent limitations: (1) the absence of reliable guidance signals for error localization, and (2) the restricted reasoning depth imposed by conventional next-token decoding paradigms. To address these issues, we propose Feedback-Triggered Regeneration (FTR), a novel framework that synergizes user feedback with enhanced decoding dynamics. Specifically, FTR activates response regeneration only upon receiving negative user feedback, thereby circumventing error propagation from faulty self-assessment while preserving originally correct outputs. Furthermore, we introduce Long-Term Multipath (LTM) decoding, which enables systematic exploration of multiple reasoning trajectories through delayed sequence evaluation, effectively overcoming the myopic decision-making characteristic of standard next-token prediction. Extensive experiments on mathematical reasoning and code generation benchmarks demonstrate that our framework achieves consistent and significant improvements over state-of-the-art prompt-based self-correction methods.

大模型自纠错多路径推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。