用单镜头演示生成上千条动作轨迹,提升机器人抓取泛化能力
One Demo is Worth a Thousand Trajectories: Action-View Augmentation for Visuomotor Policies

- 基于鱼眼相机实拍示范,生成逼真新视角图像与可行动作轨迹
- 在仿真与真实场景中,任务成功率显著提升,支持避障新环境
- 适合缺乏数据的机器人抓取研究,尤其适用于移动操作场景
用于操控的视觉-动作策略虽具潜力,但微小初始配置变化或未见障碍物易导致分布外观测,引发灾难性执行失败。本文提出一种高效数据增强框架,从单个鱼眼相机拍摄的实拍眼手示范中,生成视觉逼真的鱼眼图像序列及物理上可行的动作轨迹。引入新型高斯点云重建方法,适配广角鱼眼相机,实现含未知物体的3D场景重建与编辑;通过轨迹优化生成平滑、无碰撞、适合视图渲染的动作路径,并从新视角渲染视觉观测。仿真与真实世界实验表明,该框架在相同场景和新增障碍物需避障的场景中,均显著提升多种操控任务的成功率。
原文摘要 · Abstract (English)
Visuomotor policies for manipulation have demonstrated remarkable potential in modeling complex robotic behaviors, yet minor alterations in the robot's initial configuration and unseen obstacles easily lead to out-of-distribution observations. Without extensive data collection effort, these result in catastrophic execution failures. In this work, we introduce an effective data augmentation framework that generates visually realistic fisheye image sequences and corresponding physically feasible action trajectories from real-world eye-in-hand demonstrations, captured with a portable parallel gripper with a single fisheye camera. We introduce a novel Gaussian Splatting formulation, adapted to wide FoV fisheye cameras, to reconstruct and edit the 3D scene with unseen objects. We utilize trajectory optimization to generate smooth, collision-free, view-rendering-friendly action trajectories and render visual observations from corresponding novel views. Comprehensive experiments in simulation and the real world show that our augmentation framework improves the success rate for various manipulation tasks in both the same scene and the augmented scene with obstacles requiring collision avoidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。