arXiv:2409.14615cs.RO2024-09被引 2

让机器人同时学会选视角和操作,提升复杂场景下的灵活性。

Optimizing Active Perception for Learning Simultaneous Viewpoint Selection and Manipulation with Diffusion Policy

  • 用扩散模型结合新型视角逆运动学求解器,实现感知与操作协同优化。
  • 相比传统方法,训练效率更高,任务成功率显著提升。
  • 适合需要动态视觉的机器人应用,如手术机器人、杂乱环境操作。

机器人操作任务常依赖静态摄像头进行感知,但在机器人手术或杂乱环境等场景中,固定安装摄像头不现实。理想情况下,机器人应联合学习动态视角选择与操作策略。然而,动态视角控制需额外自由度,并与操作高度耦合,使策略学习比单臂操作更复杂。为此,我们提出一种融合扩散策略与新型‘注视’逆运动学求解器的集成学习框架,更好协调感知与操作。该框架自动优化相机朝向以选择最佳视角,同时让策略聚焦于核心操作与定位决策。实验表明,相较于直接在配置空间或末端执行器空间使用扩散策略(采用不同旋转表示),我们的方法在性能和学习效率上均表现更优。进一步分析显示,性能差异源于不同状态-动作空间中高频成分的内在差异。

原文摘要 · Abstract (English)

Robotic manipulation tasks often rely on static cameras for perception, which can limit flexibility, particularly in scenarios like robotic surgery and cluttered environments where mounting static cameras is impractical. Ideally, robots could jointly learn a policy for dynamic viewpoint and manipulation. However, dynamic viewpoint control requires additional degrees of freedom and intricate coordination with manipulation, which results in more challenging policy learning than single-arm manipulation. To address this complexity, we propose an integrated learning framework that combines diffusion policy with a novel look-at inverse kinematics solver for active perception. Our framework helps better coordinating between perception and manipulation. It automatically optimizes camera orientation for viewpoint selection, while allowing the policy to focus on essential manipulation and positioning decisions. We demonstrate that our integrated approach achieves superior performance and learning efficiency compared to directly applying diffusion policies to configuration space or end-effector space with various rotation representations. Further analysis suggests that these performance differences are driven by inherent variations in the high-frequency components across different state-action spaces.

机器人扩散模型主动感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。