用人类第一视角视频训练机器人,自动调整视角保持关键视野。
EgoAVFlow: Robot Policy Learning with Active Vision from Human Egocentric Videos via 3D Flow
- 通过3D运动流共享表示,联合学习操作与主动视觉
- 测试时动态优化视角,使任务成功率显著高于基线方法
- 无需机器人示范,适用于真实场景下复杂操作任务
第一人称人类视频为操作示范提供了可扩展的数据来源;然而,在机器人上部署这些示范需要主动控制视角以维持任务关键视野,而单纯模仿人类视角常因人类特有先验导致失效。我们提出EgoAVFlow,通过共享的3D流表示从第一人称视频中学习操作与主动视觉,支持几何可见性推理且无需机器人示范即可迁移。EgoAVFlow利用扩散模型预测机器人动作、未来3D流和相机轨迹,并在测试时通过基于预测运动与场景几何的可见性感知奖励,进行最大化奖励的去噪优化以修正视角。真实世界实验表明,在主动变化视角条件下,EgoAVFlow持续优于基于人类示范的现有方法,展示了有效可见性维护与无需机器人示范的鲁棒操作能力。
原文摘要 · Abstract (English)
Egocentric human videos provide a scalable source of manipulation demonstrations; however, deploying them on robots requires active viewpoint control to maintain task-critical visibility, which human viewpoint imitation often fails to provide due to human-specific priors. We propose EgoAVFlow, which learns manipulation and active vision from egocentric videos through a shared 3D flow representation that supports geometric visibility reasoning and transfers without robot demonstrations. EgoAVFlow uses diffusion models to predict robot actions, future 3D flow, and camera trajectories, and refines viewpoints at test time with reward-maximizing denoising under a visibility-aware reward computed from predicted motion and scene geometry. Real-world experiments under actively changing viewpoints show that EgoAVFlow consistently outperforms prior human-demo-based baselines, demonstrating effective visibility maintenance and robust manipulation without robot demonstrations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。