arXiv:2604.12126cs.AIcs.CL2026-04被引 1

用熵引导搜索提升大工具库中长程任务执行效率

Long-Horizon Plan Execution in Large Tool Spaces through Entropy-Guided Branching

  • 基于预测熵动态扩展高不确定性决策分支
  • 在SLATE基准上任务成功率显著提升,计算开销更低
  • 适合研究智能体规划与大规模工具调用的开发者

大型语言模型(LLMs)已推动工具增强型智能体的发展,使其可通过API交互实现自主推理。然而,在庞大的工具库中执行多步任务仍面临两大瓶颈:(1) 缺乏严格的计划级评估框架;(2) 由大规模工具集和长程规划带来的巨大决策空间导致的计算负担。为此,我们首先提出SLATE(面向电商的合成大规模API工具包),一个大规模、上下文感知的基准,用于自动化评估工具集成智能体。不同于静态指标,SLATE支持多样但功能有效的执行轨迹,揭示当前智能体在自我修正和搜索效率方面表现不佳。受此启发,我们进一步提出熵引导分支(EGB),一种不确定性感知的搜索算法,能动态扩展预测熵较高的决策分支。EGB优化了探索与利用的权衡,显著提升任务成功率与计算效率。在SLATE上的大量实验表明,本研究的双重贡献为构建可靠、可扩展的大型语言模型智能体提供了坚实基础。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have significantly advanced tool-augmented agents, enabling autonomous reasoning via API interactions. However, executing multi-step tasks within massive tool libraries remains challenging due to two critical bottlenecks: (1) the absence of rigorous, plan-level evaluation frameworks and (2) the computational demand of exploring vast decision spaces stemming from large toolsets and long-horizon planning. To bridge these gaps, we first introduce SLATE (Synthetic Large-scale API Toolkit for E-commerce), a large-scale context-aware benchmark designed for the automated assessment of tool-integrated agents. Unlike static metrics, SLATE accommodates diverse yet functionally valid execution trajectories, revealing that current agents struggle with self-correction and search efficiency. Motivated by these findings, we next propose Entropy-Guided Branching (EGB), an uncertainty-aware search algorithm that dynamically expands decision branches where predictive entropy is high. EGB optimizes the exploration-exploitation trade-off, significantly enhancing both task success rates and computational efficiency. Extensive experiments on SLATE demonstrate that our dual contribution provides a robust foundation for developing reliable and scalable LLM agents in tool-rich environments.

智能体规划工具调用搜索算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。