arXiv:2506.06261cs.AIcs.LG2025-06ICML被引 4

通过双重贝叶斯框架,让离线强化学习更适应不确定环境。

Reflect-then-Plan: Offline Model-Based Planning through a Doubly Bayesian Lens

  • 将规划重构为后验估计,用实时观测更新对环境动态的信念。
  • 在数据有限时仍保持稳定性能,应对高认知不确定性。
  • 适合需要安全、鲁棒策略的工业场景,如自动驾驶。

离线强化学习在在线探索代价高或不安全时至关重要,但常因数据有限导致高认知不确定性。现有方法依赖固定保守策略,限制了自适应性和泛化能力。为此,我们提出反射-规划(RefPlan),一种新颖的双重贝叶斯离线模型基础(MB)规划方法。RefPlan通过将规划重构为贝叶斯后验估计,统一了不确定性建模与MB规划。部署时,它利用实时观测更新对环境动态的信念,并通过边际化将不确定性融入MB规划。标准基准测试结果表明,RefPlan显著提升了保守离线强化学习策略的性能。尤其在高认知不确定性与数据稀缺条件下,表现出强鲁棒性,对环境动态变化也具韧性,增强了离线学习策略的灵活性、泛化性与鲁棒性。

原文摘要 · Abstract (English)

Offline reinforcement learning (RL) is crucial when online exploration is costly or unsafe but often struggles with high epistemic uncertainty due to limited data. Existing methods rely on fixed conservative policies, restricting adaptivity and generalization. To address this, we propose Reflect-then-Plan (RefPlan), a novel doubly Bayesian offline model-based (MB) planning approach. RefPlan unifies uncertainty modeling and MB planning by recasting planning as Bayesian posterior estimation. At deployment, it updates a belief over environment dynamics using real-time observations, incorporating uncertainty into MB planning via marginalization. Empirical results on standard benchmarks show that RefPlan significantly improves the performance of conservative offline RL policies. In particular, RefPlan maintains robust performance under high epistemic uncertainty and limited data, while demonstrating resilience to changing environment dynamics, improving the flexibility, generalizability, and robustness of offline-learned policies.

离线RL贝叶斯方法策略鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。