让智能体系统自动选最优模型,省钱省时还更准。
EvoRoute: Experience-Driven Self-Routing LLM Agent Systems

- 基于过往经验动态选择最合适的LLM,实时优化决策
- 在多个任务上成本降低80%,延迟减少70%以上
- 适合追求高效低成本智能体系统的开发者
由大型语言模型(LLMs)、工具和记忆模块协同的复杂智能体系统,在多轮复杂任务中展现出卓越能力。然而,其成功伴随着高昂的经济成本与严重的延迟问题,暴露出性能、成本与速度之间的关键权衡。我们将其定义为「智能体系统三难困境」:在达到顶尖性能、最小化金钱开销和确保快速完成任务之间存在固有矛盾。为破解此困局,提出EvoRoute——一种自演化模型路由范式,突破静态模型分配的局限。通过不断积累的历史经验知识库,EvoRoute在每一步动态选择帕累托最优的LLM骨干,平衡准确性、效率与资源消耗,并通过环境反馈持续优化自身选择策略。在GAIA和BrowseComp+等挑战性基准上的实验表明,将EvoRoute集成至现成智能体系统后,不仅维持或提升了性能,还实现最高80%的执行成本降低和超过70%的延迟下降。
原文摘要 · Abstract (English)
Complex agentic AI systems, powered by a coordinated ensemble of Large Language Models (LLMs), tool and memory modules, have demonstrated remarkable capabilities on intricate, multi-turn tasks. However, this success is shadowed by prohibitive economic costs and severe latency, exposing a critical, yet underexplored, trade-off. We formalize this challenge as the \textbf{Agent System Trilemma}: the inherent tension among achieving state-of-the-art performance, minimizing monetary cost, and ensuring rapid task completion. To dismantle this trilemma, we introduce EvoRoute, a self-evolving model routing paradigm that transcends static, pre-defined model assignments. Leveraging an ever-expanding knowledge base of prior experience, EvoRoute dynamically selects Pareto-optimal LLM backbones at each step, balancing accuracy, efficiency, and resource use, while continually refining its own selection policy through environment feedback. Experiments on challenging agentic benchmarks such as GAIA and BrowseComp+ demonstrate that EvoRoute, when integrated into off-the-shelf agentic systems, not only sustains or enhances system performance but also reduces execution cost by up to $80\%$ and latency by over $70\%$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。