arXiv:2410.17195cs.AIcs.CL2024-10ICLR被引 26

用预测控制提升大模型推理规划能力,减少早期错误

Non-myopic Generation of Language Models for Reasoning and Planning

  • 引入预测解码机制,基于前瞻轨迹重加权模型分布
  • 在数学、编程和智能体任务中显著提升规划准确性
  • 计算效率高于搜索基线,适合实际部署场景

大语言模型通过将复杂问题分解为序列步骤,在推理与规划任务中展现出强大能力。尽管在数学求解和编程等领域取得成功,其自回归解码的固有短视性仍导致规划不可靠、不最优。本文从最优控制视角重新审视大模型推理,提出一种新方法——预测解码(Predictive-Decoding),利用模型预测控制思想,基于前瞻轨迹对语言模型分布进行重加权,以缓解早期错误并促进非短视规划。实验表明,该方法在数学、编程及智能体任务中均有显著性能提升。此外,预测解码具有较高计算效率,在较少资源下优于传统搜索基线,为优化大模型规划能力提供了新思路。

原文摘要 · Abstract (English)

Large Language Models have demonstrated remarkable abilities in reasoning and planning by breaking down complex problems into sequential steps. Despite their success in various domains like mathematical problem-solving and coding, LLMs face challenges in ensuring reliable and optimal planning due to their inherent myopic nature of autoregressive decoding. This paper revisits LLM reasoning from an optimal-control perspective, proposing a novel method, Predictive-Decoding, that leverages Model Predictive Control to enhance planning accuracy. By re-weighting LLM distributions based on foresight trajectories, Predictive-Decoding aims to mitigate early errors and promote non-myopic planning. Our experiments show significant improvements in a wide range of tasks for math, coding, and agents. Furthermore, Predictive-Decoding demonstrates computational efficiency, outperforming search baselines with reduced computational resources. This study provides insights into optimizing LLM planning capabilities.

大模型推理规划生成预测控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。