研究企业级工具路由在规模扩大时的性能下降问题,提出有效恢复方案。
Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery

- 基于嵌入的短名单机制提升大规模工具路由准确性
- 110个智能体下路由F1下降16-23个百分点,嵌入方法恢复10-11个百分点
- 实测真实流量中恢复10-17个百分点,适合生产系统优化者
生产环境中的大语言模型助手将用户请求路由到日益增长的专业化工具库,但随着工具目录规模扩大,路由准确率如何变化?我们针对一个包含110个智能体、584个工具的企业级生产力助手的真实目录,评估了三种前沿模型在10至110个智能体规模下的单步路由表现。对于描述不充分的请求,路由F1值下降16至23个百分点。通过一个虚拟最优分析,我们将性能退化分解为‘检索差距’(模型无法召回正确工具)和‘混淆差距’(即使完美检索,最优上限仍下降10个百分点)。采用基于嵌入的短名单策略,在所有三个模型及两个提供商上,均实现全规模下+10至11个百分点的F1提升。一项生产环境标注研究(1,435条人工标注语句,三位标注者)验证了该方法在真实流量中的有效性,实际性能恢复10至17个百分点,尽管绝对性能仍比基线低10至15个百分点。
原文摘要 · Abstract (English)
Production LLM assistants route user requests to growing libraries of specialized tools, but how does routing accuracy degrade as the catalog scales? We study single-step routing on a 110-agent, 584-tool catalog from a deployed enterprise productivity assistant, evaluating three frontier models from 10 to 110 agents. Routing F1 on under-specified requests drops 16--23 percentage points across models. An oracle analysis decomposes the degradation into a \emph{retrieval} gap (the model cannot surface the right tool) and a \emph{confusion} gap (even with perfect retrieval, the oracle ceiling drops 10pp). Embedding-based shortlisting recovers +10--11pp F1 at full scale across all three models and two providers. A production annotation study (1,435 human-labeled utterances, three annotators) confirms the recovery on real traffic at +10--17pp despite 10--15pp lower absolute performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。