用多目标强化学习优化不确定感知策略,提升机器人在真实环境中的适应能力。
Domains as Objectives: Domain-Uncertainty-Aware Policy Optimization through Explicit Multi-Domain Convex Coverage Set Learning
- 将不同领域性能视为独立目标,构建多目标强化学习框架
- 通过凸覆盖集求解,使策略在随机化环境中表现更优
- 适合研究机器人控制与鲁棒强化学习的学者参考
现实世界中的不确定性是机器人任务的核心挑战,强化学习也需应对模型不确定带来的认知不确定性,这在仿真到现实的迁移中尤为明显。域随机化(DR)虽能缓解此问题,但常导致保守策略。为克服这一缺陷,引入可获取域信息的通用策略及基于循环神经网络的控制器成为新方向。本文揭示,高效优化不确定感知策略的本质可重构成多目标强化学习中的凸覆盖集(CCS)问题。通过构建以各域性能为独立目标的新马尔可夫决策过程(MDP),将不确定感知策略训练统一于多目标强化学习框架。该方法使已有多目标算法可用于域随机化训练,显著提升策略优化效率。研究聚焦线性效用函数,其与标准域随机化设定一致,并提出一系列源自多目标强化学习的算法求解凸覆盖集,验证了其对不确定感知策略性能的增强效果。
原文摘要 · Abstract (English)
The problem of uncertainty is a feature of real world robotics problems and any control framework must contend with it in order to succeed in real applications tasks. Reinforcement Learning is no different, and epistemic uncertainty arising from model uncertainty or misspecification is a challenge well captured by the sim-to-real gap. A simple solution to this issue is domain randomization (DR), which unfortunately can result in conservative agents. As a remedy to this conservativeness, the use of universal policies that take additional information about the randomized domain has risen as an alternative solution, along with recurrent neural network-based controllers. Uncertainty-aware universal policies present a particularly compelling solution able to account for system identification uncertainties during deployment. In this paper, we reveal that the challenge of efficiently optimizing uncertainty-aware policies can be fundamentally reframed as solving the convex coverage set (CCS) problem within a multi-objective reinforcement learning (MORL) context. By introducing a novel Markov decision process (MDP) framework where each domain's performance is treated as an independent objective, we unify the training of uncertainty-aware policies with MORL approaches. This connection enables the application of MORL algorithms for domain randomization (DR), allowing for more efficient policy optimization. To illustrate this, we focus on the linear utility function, which aligns with the expectation in DR formulations, and propose a series of algorithms adapted from the MORL literature to solve the CCS, demonstrating their ability to enhance the performance of uncertainty-aware policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。