arXiv:2608.21993cs.LG2026-08被引 2

提出新型上下文博弈框架,让动作可调、可过滤,提升高风险场景部署效率。

Gated Decoupled Compositional Bandits: A Unified Theory of Contextual Bandits with Supervised-Calibrated Action Scaling and Pre-Execution Gating

  • 动作由基础选择+上下文调节器组合,调节器独立训练
  • 校准的调节器能消除上下文干扰,使非平稳问题变平稳
  • 预执行门控机制支持历史数据直接初始化,无需复杂修正

我们提出广义解耦组合博弈(GDCB),一种包含三项结构创新的上下文博弈算法,超越现有线性UCB、TS、分层TS、因子化博弈、神经上下文博弈及RLHF的分类体系。在GDCB中:(i) 发送到环境的动作由离散或分层博弈选出的基础臂与上下文相关的缩放因子复合而成;(ii) 缩放因子通过独立监督学习获得,不与臂选择联合优化;(iii) 每个动作在执行前需经过预执行门控,可修改或否决该动作。我们形式化该类算法,证明四项结构性定理,并表明六类工业系统——短期租赁动态定价、临床药物剂量、信贷发放、电网需求响应、内容审核、大模型工具调用代理——均为GDCB实例,仅在组合算子、缩放族和门控机制上不同。核心成果为解耦方差缩减定理:良好校准的缩放器可消除上下文引入的方差,将非平稳问题转化为近似平稳问题。门控等价定理表明,在平稳门控下,任意先验策略收集的历史数据可作为有效冷启动初始值,无需重要性加权修正,推广了配套论文P-HITL(arXiv:2606.02595)从人类审批到任意门控的结论。在受监管、高风险领域,原本视为部署障碍的审批门控、合规规则、安全防护等,实则是实现快速部署的关键机制。配套论文在真实生产数据上验证了第一个实例(短期租赁动态定价)。

原文摘要 · Abstract (English)

We introduce Gated Decoupled Compositional Bandits (GDCB), a family of contextual bandit algorithms with three structural innovations that jointly fall outside the taxonomy of LinUCB, LinTS, HierTS, factored bandits, neural contextual bandits, and RLHF. In a GDCB system: (i) the action delivered to the environment is the composition of a nominal arm, drawn by a discrete or hierarchical bandit, with a context-dependent scaler; (ii) the scaler parameter is learned in a separate supervised loop, not jointly with arm selection; and (iii) every action passes through a pre-execution gate that may modify or veto the composed action before it reaches the environment. We formalise this class of algorithms, prove four structural theorems characterising its statistical behaviour, and show that six industrially significant systems -- short-term rental dynamic pricing, clinical drug dosing, credit origination, grid demand response, content moderation, and LLM tool-use agents -- are all instances of GDCB, differing only in the composition operator, scaler family, and gate. The central result is the Decoupling Variance Reduction theorem: a well-calibrated scaler removes context-induced variance from the arm-to-reward mapping, turning a non-stationary bandit problem into an approximately stationary one. The Gate-Induced Equivalence theorem shows that under a stationary gate, historical data collected under any prior policy is a valid warm-up initialiser without importance-sampling correction, generalising the companion P-HITL result (arXiv:2606.02595) from human approval to arbitrary gates. In regulated, high-stakes domains, constraints usually treated as deployment frictions -- approval gates, compliance rules, safety shields -- are the mechanism that makes fast deployment possible, not an obstacle to it. The companion paper validates instance 1 (STR dynamic pricing) on real production data.

上下文博弈动态定价安全控制机器学习部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。