构建首个统一的智能体推荐基准,解决复杂任务中选型难问题。
AgentSelect: Benchmark for Narrative Query-to-Agent Recommendation
- 将智能体选择转化为基于能力画像的叙事查询匹配
- 涵盖11万+查询、10万+可部署智能体和25万+交互数据
- 适合研究智能体推荐、自动化系统设计的开发者
大语言模型智能体正快速成为任务自动化的实用接口,但其生态仍缺乏系统化方法来从海量可部署配置中做出选择。现有榜单与工具/智能体基准多孤立评估组件,且在任务、指标和候选池间碎片化严重,导致关键研究空白:缺乏针对端到端智能体配置(包括基础模型与工具包组合)的查询条件监督。为此,我们提出AgentSelect,一个将智能体选择重构为基于能力画像的叙事查询-智能体推荐的基准,系统性地将异构评估数据转换为统一的仅正样本交互数据。AgentSelect整合了40多个来源的数据,包含111,179条查询、107,721个可部署智能体和251,103条交互记录,覆盖仅大模型、仅工具包及组合式智能体。分析揭示,监督模式正从密集头部复用转向长尾、近一次性监督,使基于流行度的协同过滤与图神经网络方法变得脆弱,内容感知的能力匹配成为关键。我们进一步证明,第三部分合成的组合式交互具备可学习性,在受控反事实编辑下诱导出敏感于能力的行为,并提升对真实组合的覆盖率;在AgentSelect上训练的模型还能迁移至公开智能体市场MuleRun,对未见目录带来持续增益。总体而言,AgentSelect提供了首个统一的数据与评估基础设施,为研究和加速新兴智能体生态系统奠定可复现基础。
原文摘要 · Abstract (English)
LLM agents are rapidly becoming the practical interface for task automation, yet the ecosystem lacks a principled way to choose among an exploding space of deployable configurations. Existing LLM leaderboards and tool/agent benchmarks evaluate components in isolation and remain fragmented across tasks, metrics, and candidate pools, leaving a critical research gap: there is little query-conditioned supervision for learning to recommend end-to-end agent configurations that couple a backbone model with a toolkit. We address this gap with AgentSelect, a benchmark that reframes agent selection as narrative query-to-agent recommendation over capability profiles and systematically converts heterogeneous evaluation artifacts into unified, positive-only interaction data. AgentSelectcomprises 111,179 queries, 107,721 deployable agents, and 251,103 interaction records aggregated from 40+ sources, spanning LLM-only, toolkit-only, and compositional agents. Our analyses reveal a regime shift from dense head reuse to long-tail, near one-off supervision, where popularity-based CF/GNN methods become fragile and content-aware capability matching is essential. We further show that Part~III synthesized compositional interactions are learnable, induce capability-sensitive behavior under controlled counterfactual edits, and improve coverage over realistic compositions; models trained on AgentSelect also transfer to a public agent marketplace (MuleRun), yielding consistent gains on an unseen catalog. Overall, AgentSelect provides the first unified data and evaluation infrastructure for agent recommendation, which establishes a reproducible foundation to study and accelerate the emerging agent ecosystem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。