对比四款开源路由模型在统一标准下的表现,发现选型效果更依赖候选池构成而非任务特异性。
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
- 设计统一评估协议,在四个基准上测试四款路由器性能。
- 多数路由器固定分配层级,仅vLLM Semantic Router随提示变化,但胜率未领先。
- 结果表明路由效果受候选池结构影响更大,需固定基线对照评估。
智能系统越来越多地将模型选择交由路由器处理,但开源路由器常在不同任务、候选集和执行协议下评估,难以直接比较。本文提出统一测量协议,对四种路由器实现进行混合评估,覆盖RouterBench、BFCL v4、tau2-bench和WebArena四个基准。评估290个冻结任务与2,610个锁定候选结果的组合。三个路由器输出恒定或近似恒定的层级分配;仅vLLM Semantic Router随提示内容显著变化,但在四个基准上均未取得最高成功率。Always-Mid在三个基准上与Aurelio完全一致,第四个基准误差小于0.003。对于vLLM,任务级优势检验未能发现其在任务特定性上优于按比例随机分配的盲选策略;仅在WebArena上,于协议规定的5个百分点容差内确认等价。结果表明,在此配置与控制下,观察到的性能提升更贴近所选层级构成,而非任务特异性定位。因此,固定层级基线与所选层级分布是路由器评估的必要对照;研究结论限于这些配置、候选池及冻结基准样本,并不推广至所有路由范式。
原文摘要 · Abstract (English)
Agentic systems increasingly delegate model selection to a router, yet open-source routers are usually evaluated with different tasks, candidate pools, and execution protocols, limiting direct comparison. We present a common measurement protocol and hybrid evaluation of four router implementations across RouterBench, BFCL v4, tau2-bench, and WebArena. We evaluate 290 frozen tasks against a locked matrix of 2,610 candidate outcomes. Three routers emit constant or near-constant tier assignments; only vLLM Semantic Router varies materially with prompt content, and it has the highest observed success rate on none of the four benchmarks. Always-Mid matches Aurelio exactly on three benchmarks and within 0.003 on the fourth. For vLLM, task-level superiority tests detect no task-specific advantage over a share-matched content-blind allocation; equivalence is established only on WebArena at the protocol-declared five-percentage-point margin. The results show that, under these configurations and controls, observed gains track selected-tier composition more closely than demonstrated task-specific targeting. Fixed-tier baselines and selected-tier distributions are therefore necessary controls in router evaluation; the findings are scoped to these configurations, candidate pool, and frozen benchmark samples, not to routing paradigms in general.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。