arXiv:2606.06924cs.LG2026-06被引 1

用多轮生成结果构建模型能力分布,提升大模型路由准确性

From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing

论文配图:From Sampled Outcomes to Capability Distributions: Rethinking Supervision for LLM Routing
图 1 · 摘自论文原文
  • 基于查询和输出的分布特性,构建更稳定的监督信号
  • 相比单次生成,分布感知监督使路由决策更稳定可靠
  • 适合需要精准模型选择的复杂任务场景

现有大模型路由方法通常将模型对单一查询的单次响应视为其能力标签进行训练。然而,由于大模型生成具有固有随机性,这种单次监督仅能提供噪声较大的行为观测,而非可靠的能力估计。我们发现这一假设会引入系统性噪声,降低路由策略的可靠性。为此,我们提出 DARS(分布感知路由监督)框架,从模型行为的分布视角构建监督信号。DARS 同时考虑输入侧(语义等价查询形式)与输出侧(随机生成)的不确定性,捕捉其对模型表现的影响。基于这些分布感知的观察,构建更可靠的路由监督信号。跨多种任务的实验表明,单次响应可能误导模型选择,而分布感知监督能提供更稳定的能力标签,显著改善学习到的路由行为。结果表明,可靠的大型语言模型路由应超越单次生成观测,建立在查询级模型能力分布基础上。

原文摘要 · Abstract (English)

Existing LLM routing methods typically treat a model's single response to a query as its capability label for training routers. However, because LLM generation is inherently stochastic, such single-shot supervision provides only a noisy observation of a query-model pair's behavior rather than a reliable capability estimate. We show that this assumption introduces systematic noise into routing supervision, making learned routing policies less reliable. To address this issue, we propose DARS (Distribution-Aware Routing Supervision), a framework that constructs routing supervision from a distributional view of model behavior. Instead of relying on a single generated response, DARS considers uncertainty from both the input side and the output side, capturing how semantically equivalent query formulations and stochastic generations affect model performance. Based on these distribution-aware observations, DARS builds more reliable supervision signals for routing. Experiments across diverse tasks show that single-shot labels can be misleading for model selection, while distribution-aware supervision provides more stable labels and improves learned routing behavior. Our results suggest that reliable LLM routing should move beyond single-response observations and be grounded in query-level model capability distributions.

大模型路由能力评估分布建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。