提出针对性处理不确定性框架,缓解强化学习中模型错误导致的代理误用问题。
All Models are Wrong, Knowing Where is Useful: On Model Uncertainty in Reinforcement Learning

- 基于概率模型,通过精准管理不确定性来应对模型误差
- 在真实机器人上实现高效安全探索,显著提升学习稳定性
- 适合关注模型可靠性与安全强化学习的研究者
基于模型的强化学习(MBRL)通过学习环境动态模型获取信息,有望解决机器人领域数据效率低和安全性差等难题。然而,学习到的动态模型通常存在误差,这些误差常被智能体利用,严重削弱了MBRL方法的能力。本文提出一种针对概率模型误差的不确定性处理框架,通过有目标地管理不确定性,有效缓解了模型被滥用的问题。该方法在真实硬件上实现了近期成功的学习与安全探索,并探讨了未来不确定性感知式MBRL的发展方向。
原文摘要 · Abstract (English)
Model-based reinforcement learning (MBRL) infers information about the environment from a learned dynamics model and bears the potential to address open problems such as data efficient and safe learning in robotics. However, inaccuracies of the learned dynamics model are typically exploited by the agent, substantially hampering the capabilities of MBRL methods. We present a framework for dealing with inaccuracies of probabilistic models through targeted handling of uncertainty that effectively mitigates model exploitation. We present recent successes in learning directly on hardware and safe exploration, and discuss future directions for uncertainty-aware MBRL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。