用竞价机制让大模型自报性能,降低路由成本与误差风险。
Error-Aware Reverse Auction Mechanism for Large Language Model Routing

- 模型提供商竞价时提交预测成功率和成本,由中心方统筹决策。
- 在双误差环境下仍能保持激励相容,且成本-性能权衡优于传统方法。
- 适合需要多模型高效调度的生产级大模型服务系统使用。
将每个查询路由至性价比最优的大语言模型(LLM)对平衡质量与成本至关重要,但现有路由系统依赖中心化任务平台预测模型表现,导致信息不对称与可扩展性瓶颈。本文提出基于市场的路由范式,通过反向拍卖将事前预测交给模型提供方:各提供方竞标自身预测的成功概率与执行成本。为应对提供方预测与中心评估固有的噪声,我们引入误差感知反向拍卖机制(EA-RAM),显式建模双重误差。理论证明,在双重误差下,EA-RAM具有贝叶斯激励相容性与个体理性,并给出中心理性的充分条件及明确的社会福利损失上界。实验表明,该机制在模拟与真实基准测试中均对双重误差具有鲁棒性,实现更优的成本-性能帕累托前沿,且当提供方贡献本地信息时收益进一步提升,验证了其实际有效性。
原文摘要 · Abstract (English)
Routing each query to a cost-effective large language model (LLM) is critical for balancing quality and cost, yet most routers rely on a centralized task center to predict model performance, creating an information-risk mismatch and a scalability bottleneck as the model pool grows. We propose a market-based routing paradigm that shifts ex-ante prediction to LLM providers via a reverse auction, where providers bid with self-predicted success probabilities and execution costs. To account for inherently noisy provider predictions and center evaluations, we introduce the \textit{\textbf{E}rror-\textbf{A}ware \textbf{R}everse \textbf{A}uction \textbf{M}echanism} (EA-RAM), which explicitly models this inherent Dual Error. We prove that EA-RAM is Bayesian incentive compatible and individually rational under the Dual Error, establish sufficient conditions for center rationality, and derive an explicit welfare-loss bound. We further identify robustness effects: opposite-signed errors can cancel, vanishing-tail link functions (e.g., logistic) stabilize clear-cut cases via saturation, and extra noise smooths belief maps, reducing the gains from marginal manipulation. Experiments on simulations and real-world benchmarks show that EA-RAM is robust to the Dual Error and achieves a better cost--performance Pareto frontier than centralized baselines, with additional gains when providers contribute local information, validating its practical effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。