让机器人在移动操作中自主评估风险,实时调整决策谨慎程度。
Risk-Aware Reinforcement Learning for Mobile Manipulation
- 用分布强化学习训练教师策略,通过风险度量调节行为谨慎性。
- 在未建图环境中实现反应式运动,最差情况表现显著提升。
- 可通过模仿学习将风险感知能力迁移到依赖深度观测的视觉电机策略。
为使机器人从实验室走向日常环境,必须能推理自身动作的风险并做出知情的、风险敏感的决策。尤其对于需在动态非结构化空间中同时导航与交互的移动操作任务,现有全身控制器通常缺乏显式的不确定性下风险敏感决策机制。本文首次(i)学习基于自参考深度观测的风险感知视觉电机策略,支持运行时可调风险敏感度;(ii)证明风险感知行为可通过模仿学习(IL)迁移至基于自参考深度观测的视觉电机策略。方法首先利用分布强化学习(DRL)训练特权教师策略,采用风险中性分布批评者;随后对批评者的回报分布施加畸变风险度量,生成风险调整后的优势估计用于策略更新,以实现多种风险感知行为。再通过模仿学习将教师策略蒸馏为条件于自参考深度观测的风险感知学生策略。大量实验表明,所训练的视觉电机策略在未建图环境中执行反应式全身运动时,能展现出风险感知行为(具体表现为更优的最差情况性能),并利用实时深度观测进行感知。
原文摘要 · Abstract (English)
For robots to successfully transition from lab settings to everyday environments, they must begin to reason about the risks associated with their actions and make informed, risk-aware decisions. This is particularly true for robots performing mobile manipulation tasks, which involve both interacting with and navigating within dynamic, unstructured spaces. However, existing whole-body controllers for mobile manipulators typically lack explicit mechanisms for risk-sensitive decision-making under uncertainty. To our knowledge, we are the first to (i) learn risk-aware visuomotor policies for mobile manipulation conditioned on egocentric depth observations with runtime-adjustable risk sensitivity, and (ii) show risk-aware behaviours can be transferred through Imitation Learning (IL) to a visuomotor policy conditioned on egocentric depth observations. Our method achieves this by first training a privileged teacher policy using Distributional Reinforcement Learning (DRL), with a risk-neutral distributional critic. Distortion risk-metrics are then applied to the critic's predicted return distribution to calculate risk-adjusted advantage estimates used in policy updates to achieve a range of risk-aware behaviours. We then distil teacher policies with IL to obtain risk-aware student policies conditioned on egocentric depth observations. We perform extensive evaluations demonstrating that our trained visuomotor policies exhibit risk-aware behaviour (specifically achieving better worst-case performance) while performing reactive whole-body motions in unmapped environments, leveraging live depth observations for perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。