用结果导向约束提升离线强化学习安全性与泛化能力
Beyond Non-Expert Demonstrations: Outcome-Driven Action Constraint for Offline Reinforcement Learning
- 基于动作结果是否安全来评估,而非依赖原始数据动作分布
- 在MuJoCo和迷宫任务中实现更优轨迹拼接与未见状态适应
- 适合处理真实场景中的非专家行为数据,提升学习鲁棒性
针对使用现实世界非专家数据进行离线强化学习的挑战,本文提出一种名为结果驱动动作灵活性(ODAF)的新方法。该方法旨在降低对行为策略经验动作分布的依赖,从而减轻不良示范带来的负面影响。具体而言,引入一种保守奖励机制,通过判断动作结果是否满足安全要求(即是否处于状态支持区域内)来评估动作,而非仅依据其在离线数据中的出现概率。在广泛使用的MuJoCo和多种迷宫基准上的实验证明,结合不确定性量化技术的ODAF方法,能有效容忍未见转移,改善轨迹拼接效果,并显著提升智能体从真实非专家数据中学习的能力。
原文摘要 · Abstract (English)
We address the challenge of offline reinforcement learning using realistic data, specifically non-expert data collected through sub-optimal behavior policies. Under such circumstance, the learned policy must be safe enough to manage distribution shift while maintaining sufficient flexibility to deal with non-expert (bad) demonstrations from offline data.To tackle this issue, we introduce a novel method called Outcome-Driven Action Flexibility (ODAF), which seeks to reduce reliance on the empirical action distribution of the behavior policy, hence reducing the negative impact of those bad demonstrations.To be specific, a new conservative reward mechanism is developed to deal with distribution shift by evaluating actions according to whether their outcomes meet safety requirements - remaining within the state support area, rather than solely depending on the actions' likelihood based on offline data.Besides theoretical justification, we provide empirical evidence on widely used MuJoCo and various maze benchmarks, demonstrating that our ODAF method, implemented using uncertainty quantification techniques, effectively tolerates unseen transitions for improved "trajectory stitching," while enhancing the agent's ability to learn from realistic non-expert data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。