用单次真实演示生成逼真物理一致的机械物体操作数据,提升机器人泛化能力。
AOMGen: Photoreal, Physics-Consistent Demonstration Generation for Articulated Object Manipulation
- 从一次真实扫描和演示出发,合成多视角、同步动作与状态数据。
- 微调后模型在未见物体和布局上成功率从0%提升至88.7%。
- 适合需要高保真物理仿真数据的机器人操控研究者。
视觉-语言-动作(VLA)与世界模型方法虽提升了机器人操作的泛化能力,但其成功依赖大量昂贵的真实示范数据,尤其在关节类物体精细操作中更为显著。为此,我们提出AOMGen——一种可扩展的关节物体操作数据生成框架。该框架基于单次真实扫描、一次示范及现成数字资产库,生成具有验证物理状态的逼真训练数据。框架同步生成多视角RGB序列,时间对齐动作指令与关节状态、接触标注,并系统变换相机视角、物体风格与姿态,将单一执行扩展为多样化数据集。实验表明,在AOMGen数据上微调的VLA策略,使未见过物体与布局上的成功率从0%提升至88.7%。
原文摘要 · Abstract (English)
Recent advances in Vision-Language-Action (VLA) and world-model methods have improved generalization in tasks such as robotic manipulation and object interaction. However, Successful execution of such tasks depends on large, costly collections of real demonstrations, especially for fine-grained manipulation of articulated objects. To address this, we present AOMGen, a scalable data generation framework for articulated manipulation which is instantiated from a single real scan, demonstration and a library of readily available digital assets, yielding photoreal training data with verified physical states. The framework synthesizes synchronized multi-view RGB temporally aligned with action commands and state annotations for joints and contacts, and systematically varies camera viewpoints, object styles, and object poses to expand a single execution into a diverse corpus. Experimental results demonstrate that fine-tuning VLA policies on AOMGen data increases the success rate from 0% to 88.7%, and the policies are tested on unseen objects and layouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。