arXiv:2412.10827cs.CLcs.AI2024-12ICML被引 12

用自训练视角重构思维链,提升大模型推理效率与准确性

Rethinking Chain-of-Thought from the Perspective of Self-Training

  • 将思维链与自训练结合,动态优化推理过程
  • 在多个任务上实现性能提升,计算开销更低
  • 适合需要高效精准推理的应用场景

思维链(CoT)推理已成为激活大语言模型潜在能力的有效方法。我们观察到,CoT推理与自训练共享核心目标:通过迭代利用模型生成的信息,逐步降低预测不确定性。基于此洞察,我们提出一种新型CoT框架以提升推理性能。该框架包含两个关键组件:(i) 任务特定提示模块,优化初始推理过程;(ii) 自适应推理迭代模块,动态调整推理流程,解决先前CoT方法中存在的过度推理及连续推理步骤高度相似的问题。大量实验表明,所提方法在性能与计算效率方面均具有显著优势。

原文摘要 · Abstract (English)

Chain-of-thought (CoT) reasoning has emerged as an effective approach for activating latent capabilities in LLMs. Interestingly, we observe that both CoT reasoning and self-training share the core objective: iteratively leveraging model-generated information to progressively reduce prediction uncertainty. Building on this insight, we propose a novel CoT framework to improve reasoning performance. Our framework integrates two key components: (i) a task-specific prompt module that optimizes the initial reasoning process, and (ii) an adaptive reasoning iteration module that dynamically refines the reasoning process and addresses the limitations of previous CoT approaches, \ie over-reasoning and high similarity between consecutive reasoning iterations. Extensive experiments demonstrate that the proposed method achieves significant advantages in both performance and computational efficiency.

思维链自训练推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。