提出可验证的评估方法,揭示多模型路由中真实可用的性能提升有限。
Opportunity Is Not Realizability: Selection-Valid Diagnostics for Multi-LLM Routing
- 分离三种可评估指标,避免选择偏差影响结果
- 实测显示部署级路由仅实现14.4%的理论提升空间
- 证明了强路由仍优于单模型,但差距仍有余地
Oracle路由评估一个语言模型池在按查询选择时能获得多少收益,但现有方法存在两个缺陷:一是以相同样本上选出的最佳固定模型为基准,破坏了配对推理;二是全信息oracle能看到部署路由器无法观察到的结果。本文区分三种可估量目标(结果型oracle机会、基于预答信号的贝叶斯最优增益、学习路由的保留集增益),并证明在选择最佳固定模型或路由家族成员时依然有效的置信区间。通过信号信息夹逼法与(1-1/e)贪心保证,可在子模互补覆盖下构建紧凑模型池。在六个模型族、八个检查点、四个基准上的实验表明,所有任务的群体oracle差距为9.7–30.7分,而最强的可部署提示路由仅恢复7.5–14.4%,且十一种策略中最佳者的置信区间下限始终为零。真实可实现的oracle机会很小且可证。
原文摘要 · Abstract (English)
Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes. We separate three estimands (outcome-oracle opportunity, the Bayes-optimal gain from a declared pre-answer signal, and the held-out gain of a learned router) and prove selection-valid confidence intervals that survive choosing the best fixed model or the best member of a router family, a signal-information sandwich, and a $(1-1/e)$ greedy guarantee for building compact pools from submodular complementary coverage. On eight checkpoints from six families over four benchmarks, selection-valid intervals certify a population oracle gap of $9.7$--$30.7$ points on every task, yet the strongest deployable prompt router recovers only $7.5$--$14.4\%$ of it, and the simultaneous interval for the best of eleven tested policies has lower limit zero throughout. The realizable share of oracle opportunity is small and certifiable: strong routers beat the best fixed model, and most of the gap remains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。