arXiv:2605.29788cs.AIcs.LG2026-05

提出可认证的嵌套因果多阶段决策方法,实现安全渐进部署。

Certified Policy Optimisation for Nested Causal Bandits via PAC-Bayes Risk

  • 基于因果结构分解信念,递归执行策略以应对多时标依赖。
  • 仅用历史数据即可提供可证伪的部署风险评估,随数据积累风险下降。
  • 适合需要安全升级旧系统、追求可解释性与风险可控的决策场景。

关键的序列决策往往非单一时标:战略决策会因果性地影响后续战术选择的上下文;标准强化学习与带宽理论无法捕捉时标间的因果耦合。本文将此类问题形式化为嵌套上下文因果带宽(NCCBs),一种分层结构因果模型(SCM),其中每一层的动作决定下一层的上下文分布。我们提出嵌套因果汤普森采样(NCTS),每轮抽取一个机制分解后的信念,并在其上递归行动。主要理论成果是首个因果 PAC-Bayes 超额风险界,可仅凭历史数据对任意候选策略进行离线、非监督、任意时间点的风险认证,回答部署问题:能否信任该智能体在此处?实验在分层因果模型上验证:相比同函数类上的联合回归(RFF-GP),因子化机制后验在外部分布漂移下零样本迁移表现显著更优;递归元到内层决策优于联合承诺方案;证书风险随离线数据积累显著收缩。结合上述结果,建立渐进式可认证移交机制:每个时标在收益可被认证时独立从旧控制器切换至 NCTS。

原文摘要 · Abstract (English)

Critical sequential decisions are rarely single-timescale: a strategic decision causally shapes the context in which every subsequent tactical choice is made; standard bandit and reinforcement-learning theory does not capture this causal coupling between timescales. We formalise the problem class as Nested Contextual Causal Bandits (NCCBs), a hierarchical SCM where each level's action sets the next level's context distribution, and propose Nested Causal Thompson Sampling (NCTS), which draws one mechanism-factorised belief per episode and acts recursively under it. Our main theoretical result is a causal PAC-Bayesian excess-risk bound that certifies any candidate deployment policy from historic data alone, off-policy and anytime, answering the deployment question: can we trust this agent here, and at what risk? Experiments on a hierarchical SCM show that, against a matched RFF-GP joint regression on the same function class, the factorised SCM-mechanism posterior transfers significantly better zero-shot under exogenous distribution shifts, the recursive meta-to-inner commit significantly dominates the joint-commit alternative in distribution, and the certificate significantly contracts as offline data accumulates. Combining these results, we establish progressive certified handover, a safe-deployment method: each timescale flips from a legacy controller to NCTS when gains can be certified, independently of the others.

因果推断决策优化安全部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。