通过动态物体运动增强数据多样性,提升机器人操作的空间泛化能力
MOVE: A Simple Motion-Based Data Collection Paradigm for Spatial Generalization in Robotic Manipulation
- 在演示中主动移动环境中的可动物体,生成多样化空间配置
- 仿真环境下成功率达39.1%,比静态数据提升76.1%
- 显著提高数据效率,部分任务达2–5倍增益,适合强化学习数据收集
模仿学习在机器人操作中展现巨大潜力,但实际部署受限于数据稀缺。尽管已有大规模数据集研究,但在空间泛化方面仍存在明显差距。我们发现关键瓶颈在于:每条轨迹通常在单一静态空间配置下采集,包括固定物体位置、目标位置和相机视角,严重限制了空间信息的多样性。为此,我们提出运动增强数据采集范式(MOVE),通过在每次演示中对环境中可动物体注入运动,隐式生成密集且多样的空间配置。该方法在仿真与真实环境中的广泛实验验证其有效性。例如,在需强空间泛化能力的任务中,MOVE实现平均成功率39.1%,相较静态数据采集方式(22.2%)提升76.1%,并在某些任务上实现2–5倍的数据效率增益。
原文摘要 · Abstract (English)
Imitation learning method has shown immense promise for robotic manipulation, yet its practical deployment is fundamentally constrained by the data scarcity. Despite prior work on collecting large-scale datasets, there still remains a significant gap to robust spatial generalization. We identify a key limitation: individual trajectories, regardless of their length, are typically collected from a \emph{single, static spatial configuration} of the environment. This includes fixed object and target spatial positions as well as unchanging camera viewpoints, which significantly restricts the diversity of spatial information available for learning. To address this critical bottleneck in data efficiency, we propose \textbf{MOtion-Based Variability Enhancement} (\emph{MOVE}), a simple yet effective data collection paradigm that enables the acquisition of richer spatial information from dynamic demonstrations. Our core contribution is an augmentation strategy that injects motion into any movable objects within the environment for each demonstration. This process implicitly generates a dense and diverse set of spatial configurations within a single trajectory. We conduct extensive experiments in both simulation and real-world environments to validate our approach. For example, in simulation tasks requiring strong spatial generalization, \emph{MOVE} achieves an average success rate of 39.1\%, a 76.1\% relative improvement over the static data collection paradigm (22.2\%), and yields up to 2--5$\times$ gains in data efficiency on certain tasks. Our code is available at https://github.com/lucywang720/MOVE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。