arXiv:2606.19376cs.LGcs.AI2026-06

在用户反馈有限下实现低成本且符合服务质量承诺的LLM请求调度

Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees

论文配图:Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees
图 1 · 摘自论文原文
  • 基于稀疏单向用户反馈在线学习成本最优路由策略
  • 实验显示可降低2.2倍运营成本,且无需针对不同任务调参
  • 适合需要严格服务保障的商业LLM部署场景

大型语言模型(LLM)应用的推理成本因需求激增和基础设施费用上升而快速攀升。用户期望高质量响应,商业场景中这由服务水平协议(SLA)正式规定,导致成本与质量之间存在根本矛盾。现有成本感知的LLM请求调度方法虽有潜力缓解此矛盾,但依赖完整反馈信号、离线训练、大量工作负载调参,且多数缺乏SLA保证或运行时自适应能力。我们提出SLARouter,一种从生产系统中稀疏、单向用户反馈中在线学习成本最优策略的路由算法。SLARouter在成本最优性和严格遵守SLA方面提供理论保障。在多种LLM基准测试上的实验表明,SLARouter无需针对各基准进行调参即可满足SLA约束,相比现有基线最多降低2.2倍运营成本。

原文摘要 · Abstract (English)

Inference costs for large language model (LLM) applications are rapidly growing, driven by surging demand and rising infrastructure cost. Users expect high-quality responses, and in commercial settings this is formally codified in Service Level Agreements (SLAs), creating a fundamental tension between cost and quality. Recent progress on cost-aware LLM request routing has shown potential to resolve this tension, but existing approaches rely on complete feedback signals, offline training, extensive per-workload tuning, and most lack SLA guarantees or inference-time adaptivity. We introduce SLARouter, an online routing algorithm that learns a cost-optimal policy from the sparse, one-sided user feedback available in production systems. SLARouter provides theoretical guarantees for both cost optimality and strict SLA compliance. Experiments across a wide range of LLM benchmarks show that SLARouter satisfies SLA constraints without the need for per-benchmark tuning, reducing operating cost by up to 2.2x over existing baselines.

LLM调度成本优化SLA保障

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。