arXiv:2607.03436cs.LG2026-07被引 2

揭示大模型路由差距中噪声与真实优势的构成比例

How Much of the Routing Gap Is Real? Decomposing the Router-to-Oracle Gap into Reproducible Specialist Advantage and Single-Draw Label Noise

  • 将路由误差分解为可复现优势和单次采样噪声
  • 实测显示噪声占差距12%~36%,最难任务中接近一半
  • 提供可复用的多抽样协议,适配评测基准优化

在真实开放模型池中,报告的路由器与理想最优器之间的差距中,12%至36%源于单次采样标签噪声,任何单一提交的路由器都无法捕捉;而多数差距是真实的、可恢复的专业模型优势。本文证明了这种可恢复性不对称性,并发布一种测量协议。由于随机解码,理想最优器本质是单次伯努利抽样,非可复现属性。我们将期望最优器分解为可复现部分 $O^{ ext{repro}}$ 和不可恢复的单次选择底限 $Δ$。主结论是:该底限无法被任何单次提交的路由器(确定或随机)覆盖,但可通过测试时采样(如 best-of-$K$)在相同预算下超越独立池的单次抽样最优器。该上限不依赖跨模型独立性,真正指向单次提交的选择限制,而非信息缺失。通过 LLMRouterBench(33个模型,391,645个实例)构建每查询的 $T=0.2$ 单次抽样联合最优器,其20分差距即由随机抽样构成。因 $k=1$ 下 $O^{ ext{repro}}$ 不可识别,采用单侧、去相关边界对 $k\geq20$ 重新估计。在三个受控开放模型重生成任务(算术、竞赛数学、非数学科学)中,单次噪声为差距的重要组成部分,在未饱和任务中更大,最难查询中占比接近一半。论文发布多抽样最优器协议,供路由评测采纳。

原文摘要 · Abstract (English)

On real open-model pools, 12--36% of the reported router-to-oracle gap is single-draw label noise that no single-commit router can capture, while the majority is genuine, recoverable specialist advantage; this work proves why (a recoverability asymmetry) and releases a protocol to measure it. Routing among large language models (LLMs) trades cost for quality, motivated by the gap between learned routers and a per-instance oracle. But under stochastic decoding that oracle is a single Bernoulli draw, not a reproducible property. We recast the question structurally: the expected oracle decomposes as $O^{\exp}=O^{\mathrm{repro}}+Δ$, into reproducible single-commit headroom $O^{\mathrm{repro}}$ and a non-negative single-commit selection floor $Δ$. Our main result is a recoverability asymmetry: this floor is closed by no single-commit router (deterministic or randomized), yet is provably recovered by test-time sampling: best-of-$K$ on the committed model, at the oracle's own budget, dominates the independent-pool single-draw oracle. This cap needs no cross-model independence, pinning "not recoverable" to single-commit selection, not to information. The floor's magnitude is a prospective, conservative localization, not an audit: LLMRouterBench (33 models, 391,645 instances) builds its oracle as a per-query union of single $T=0.2$ draws, so its 20-point gap is by construction a union of stochastic draws; since $O^{\mathrm{repro}}$ is non-identifiable at $k=1$, we re-estimate by fresh $k\ge20$ resampling under one-sided, dependence-corrected bounds. Across three controlled open-model re-generations (arithmetic, competition math, and non-math science), single-draw noise is a substantial minority of the gap, larger on unsaturated benchmarks and approaching half on the hardest queries. We release a multi-sample oracle protocol that routing benchmarks can adopt.

大模型路由评测基准随机性分析可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。