arXiv:2512.01188cs.ROcs.AI2025-12NeurIPS被引 6

让机器人学会主动探查信息,提升复杂任务表现。

Real-World Reinforcement Learning of Active Perception Behaviors

  • 利用训练时的额外传感器辅助学习,高效生成主动感知策略。
  • 仅需少量演示和粗糙初始策略,即可在8个任务中超越已有方法。
  • 适合需要在信息不全情况下自主决策的机器人应用。

机器人的即时感官观测往往无法揭示任务相关状态信息。在部分可观测条件下,最优行为通常需主动采取动作以获取缺失信息。当前主流机器人学习方法难以生成此类主动感知行为。本文提出一种简单有效的现实世界机器人学习方案——非对称优势加权回归(AAWR),利用训练时可访问的“特权”传感器,构建高质量的特权价值函数,辅助估计目标策略的优势。基于少量潜在次优示范和易得的粗略策略初始化,AAWR能快速习得主动感知行为并显著提升任务性能。在3种机器人、8个操作任务上的评估显示,该方法生成了可靠主动感知行为,全面优于此前所有方法。当初始化为在主动感知任务中表现不佳的‘通用’机器人策略时,AAWR仍能高效生成信息收集行为,使其在严重部分可观测条件下完成操作任务。

原文摘要 · Abstract (English)

A robot's instantaneous sensory observations do not always reveal task-relevant state information. Under such partial observability, optimal behavior typically involves explicitly acting to gain the missing information. Today's standard robot learning techniques struggle to produce such active perception behaviors. We propose a simple real-world robot learning recipe to efficiently train active perception policies. Our approach, asymmetric advantage weighted regression (AAWR), exploits access to "privileged" extra sensors at training time. The privileged sensors enable training high-quality privileged value functions that aid in estimating the advantage of the target policy. Bootstrapping from a small number of potentially suboptimal demonstrations and an easy-to-obtain coarse policy initialization, AAWR quickly acquires active perception behaviors and boosts task performance. In evaluations on 8 manipulation tasks on 3 robots spanning varying degrees of partial observability, AAWR synthesizes reliable active perception behaviors that outperform all prior approaches. When initialized with a "generalist" robot policy that struggles with active perception tasks, AAWR efficiently generates information-gathering behaviors that allow it to operate under severe partial observability for manipulation tasks. Website: https://penn-pal-lab.github.io/aawr/

强化学习主动感知机器人学习部分可观测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。