让机器人在看与动之间协同决策,提升遮挡环境下的操作能力
Learning to See While Learning to Act: Diffusion Models for Active Perception in Robot Imitation

- 用扩散模型联合优化动作与主动视角选择
- 在RLBench上性能比之前方法最高提升34%
- 零样本迁移到真实世界,处理严重遮挡
多数模仿学习方法假设桌面场景中物体完全可见。实际中物体常被遮挡,需机器人同时进行搜索与操作,从有限示范中学习这种耦合行为仍具挑战。本文提出See2Act,通过将动作去噪与视角优化耦合,在测试时基于主动推断的视角序列预测动作。策略利用离线示范中关键帧动作锚定的相机位姿进行训练,隐式学习‘何时看’的同时学会‘如何操作’。实验证明,在Ravens中该策略能恢复有信息量的视角;在RLBench任务上性能相比先前方法最高提升34%。在真实世界中,我们通过数字孪生收集50组示范,仅用深度观测即实现零样本仿真到现实迁移,成功完成拾取放置任务,证明所学视角推理可有效应对部分可观测性。
原文摘要 · Abstract (English)
Most imitation learning methods assume full observability in table-top settings. In practice, objects are often occluded, requiring robots to both search and act, and learning this coupled behavior from limited demonstrations remains challenging. We propose See2Act, an imitation learning approach that conditions action prediction on a sequence of actively-inferred viewpoints at test time, by coupling action denoising with viewpoint refinement. The policy is trained using camera poses anchored to keyframe actions from offline demonstrations, enabling implicit learning of where to see, while learning how to act. We empirically demonstrate that in Ravens the policy recovers informative viewpoints under severe occlusions, and on RLBench tasks it improves performance by up to 34% over prior methods. In the real world, we collect 50 demonstrations in a digital twin and achieve zero-shot sim-to-real transfer on pick-and-place tasks using depth observations. The policy handles significant occlusions, showing that learned viewpoint reasoning enables robust manipulation under partial observability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。