解决视觉数据中的隐性混淆问题,提升离线强化学习鲁棒性
Causal Flow Q-Learning for Robust Offline Reinforcement Learning
- 从因果视角设计新目标,应对观测混淆带来的偏差
- 在25个像素任务上成功率比现有方法高出120%
- 适合处理感官能力不匹配的离线训练场景
近期基于流匹配的表达性强策略在强化学习中表现优异,因其能从离线数据中建模复杂动作分布。这些算法依赖标准策略梯度,假设数据中无未测量混杂因素。然而,在演示者与学习者感官能力不一致时,像素级示范数据可能产生隐性混杂偏差。本文从因果角度分析离线强化学习中的混杂观测问题,提出一种新的因果离线强化学习目标,优化因混杂偏差可能导致的最差性能。基于此目标,我们设计了一种实用实现,通过深度判别器评估目标策略与名义行为策略间的差异,从混杂示范数据中学习表达性强的流匹配策略。在25个像素任务上的实验表明,所提出的抗混杂增强方法成功率比现有最先进的无意识混杂方法高出120%。
原文摘要 · Abstract (English)
Expressive policies based on flow-matching have been successfully applied in reinforcement learning (RL) more recently due to their ability to model complex action distributions from offline data. These algorithms build on standard policy gradients, which assume that there is no unmeasured confounding in the data. However, this condition does not necessarily hold for pixel-based demonstrations when a mismatch exists between the demonstrator's and the learner's sensory capabilities, leading to implicit confounding biases in offline data. We address the challenge by investigating the problem of confounded observations in offline RL from a causal perspective. We develop a novel causal offline RL objective that optimizes policies' worst-case performance that may arise due to confounding biases. Based on this new objective, we introduce a practical implementation that learns expressive flow-matching policies from confounded demonstrations, employing a deep discriminator to assess the discrepancy between the target policy and the nominal behavioral policy. Experiments across 25 pixel-based tasks demonstrate that our proposed confounding-robust augmentation procedure achieves a success rate 120\% that of confounding-unaware, state-of-the-art offline RL methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。