从人类示范中学习双手机器人主动感知策略,提升复杂操作鲁棒性。
Vision in Action: Learning Active Perception from Human Demonstrations
- 通过虚拟现实接口捕捉人类主动视觉行为,实现人机共享观察空间。
- 在三个多阶段任务中表现优于基线系统,尤其在视觉遮挡场景下。
- 采用6自由度机械颈部与异步3D场景更新,减少操作延迟和眩晕感。
我们提出视觉在行动(Vision in Action, ViA),一种用于双手机器人操作的主动感知系统。ViA 直接从人类示范中学习任务相关的主动感知策略(如搜索、追踪和聚焦)。硬件上,系统采用一个简单但有效的6-DoF机械颈部,实现类人头动。为捕捉人类主动感知策略,设计了一种基于虚拟现实(VR)的遥操作界面,建立机器人与操作者之间的共享观察空间。为缓解因机器人物理动作延迟导致的VR晕动症,该界面使用中间3D场景表示,使操作端能实时渲染视角,同时异步用机器人最新观测更新场景。这些设计共同支持在三个涉及视觉遮挡的复杂多阶段双手操作任务中,学习到鲁棒的视觉运动策略,显著优于基线系统。
原文摘要 · Abstract (English)
We present Vision in Action (ViA), an active perception system for bimanual robot manipulation. ViA learns task-relevant active perceptual strategies (e.g., searching, tracking, and focusing) directly from human demonstrations. On the hardware side, ViA employs a simple yet effective 6-DoF robotic neck to enable flexible, human-like head movements. To capture human active perception strategies, we design a VR-based teleoperation interface that creates a shared observation space between the robot and the human operator. To mitigate VR motion sickness caused by latency in the robot's physical movements, the interface uses an intermediate 3D scene representation, enabling real-time view rendering on the operator side while asynchronously updating the scene with the robot's latest observations. Together, these design elements enable the learning of robust visuomotor policies for three complex, multi-stage bimanual manipulation tasks involving visual occlusions, significantly outperforming baseline systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。