利用代理奖励提升大模型路由的样本效率与精度平衡
Correlation-Aware Contextual Bandits with Surrogate Rewards for LLM Routing

- 融合真实反馈与机器学习生成的代理奖励,应对臂间相关性
- 在代理奖励可靠时加速学习,最坏情况下仍保持理论保障
- 适合需要高效调度大模型的推荐系统与服务架构
我们研究具有相关臂和可访问机器学习模型生成的代理奖励信号的上下文关联老虎机问题,动机来自大语言模型(LLM)路由等应用。与传统仅依赖老虎机反馈并假设臂间条件独立的设定不同,本设置允许上下文相关的臂间相关性以及可能噪声或误设的辅助奖励信息。我们提出两种互补算法:耦合奖励混合方法在代理信号可靠时融合真实与代理奖励以加速学习;解耦预测混合方法分别维护老虎机反馈与代理奖励的估计器,并自适应组合其预测。该解耦设计对代理误设具有鲁棒性,在最坏情况下恢复与纯奖励老虎机方法相当的遗憾保证,同时在代理预测足够有效时实现更优遗憾。我们为两种方法提供了理论遗憾分析,并在不同准确率-成本权衡下对LLM路由基准进行评估。结果表明,相比标准上下文关联老虎机基线和强静态路由方法,新方法显著提升样本效率,并持续获得更优的准确率-成本权衡。
原文摘要 · Abstract (English)
We study contextual bandit problems with correlated arms and access to surrogate reward signals produced by a machine learning model, motivated by applications such as large language model (LLM) routing. Unlike classical contextual bandits that rely solely on bandit feedback and assume conditional independence across arms, our setting allows context-dependent inter-arm correlations and auxiliary reward information that may be noisy or misspecified. We propose algorithms that leverage such surrogate rewards through two complementary designs. A coupled reward-mixing approach pools true and surrogate rewards to accelerate learning when surrogate signals are reliable, while a decoupled prediction-mixing approach maintains separate estimators for bandit feedback and surrogate rewards and adaptively combines their predictions. This decoupling yields robustness to surrogate misspecification, recovering regret guarantees comparable to reward-only bandit methods in the worst case, while achieving improved regret when surrogate predictions are sufficiently informative. We provide theoretical regret analyses for both approaches and evaluate them on LLM routing benchmarks under varying accuracy versus cost trade-offs. The results demonstrate improved sample efficiency and consistently better accuracy-cost trade-offs compared to standard contextual bandit baselines and strong static routing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。