arXiv:2512.03057stat.MLcs.AI2025-12被引 3

提出可保证推理效率的智能模型路由方法,避免盲目调用专家模型。

A note on conditional PAC-efficient reasoning in large language model routing

  • 基于预设条件集设计受限的条件化路由策略
  • 在有限样本下实现可靠路由,接近最优专家使用率
  • 适合对推理效率与可靠性有严格要求的应用场景

我们研究大语言模型推理中的无分布风险控制问题。形式化了满足大概率近似正确(PAC)保证的逐点条件效率,并证明其导致近乎不可能的路由器:当快速模型损失超过目标时,几乎在所有输入上都必须以至少1减去预定误差水平的概率路由到专家模型。因此,我们提出基于预设条件集族的受限条件化公式,并设计了一种显式路由器。该路由器在有限样本下实现条件有效性,且在分离性和间隔条件下达到近似最优的专家使用率。核心洞察是:条件化程度决定了无分布可靠性与计算节省能否共存——逐点控制过强不可行,而结构化的集合控制仍可实现。

原文摘要 · Abstract (English)

We study distribution-free risk control for model routing, motivated by large language model reasoning. We formalize pointwise conditional efficiency under a probably approximately correct guarantee and show that it forces a nearly impossible router: at almost every input where the fast model exceeds the target loss, the algorithm must route to the expert with probability at least one minus the prescribed error level. We therefore introduce a restricted conditional formulation based on a prespecified family of conditioning sets, together with an explicit router. The proposed router achieves finite-sample conditional validity and, under separation and margin conditions, near-oracle expert usage. The main insight is that the level of conditioning determines whether distribution-free reliability can coexist with computational savings: pointwise control is too strong, whereas structured setwise control remains feasible.

模型路由PAC学习大模型推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。