用生成视频提升机器人抓取的泛化能力,真实场景表现提升92%以上。
EMMA: Generalizing Real-World Robot Manipulation via Generative Visual Transfer
- 通过扩散模型生成多视角一致的机器人操作视频,支持视觉编辑且保持3D结构
- 在1800次模拟与真实测试中,零样本视觉下性能比纯真实数据训练高92%
- 自适应混合训练策略增强难点样本学习,适合研究机器人视觉-语言-动作泛化
VLA模型的泛化依赖多样化的训练数据,但获取跨物体外观的大规模机器人操作数据成本高昂。为此,我们提出具身操作媒体适配框架EMMA,结合生成数据引擎与高效训练流程。提出DreamTransfer——一种基于扩散Transformer的架构,可生成多视角一致、几何合理的具身操作视频,支持通过提示修改前景、背景和光照,同时保持3D结构与几何有效性。采用真实与生成数据混合训练集,提出AdaMix策略,根据策略表现自适应加权样本,强化困难样本学习。全面评估表明,DreamTransfer生成视频在多视角一致性、几何准确性和文本条件精度上显著优于现有方法。在模拟与真实机器人环境中共开展1800余次测试。在零样本视觉设置下,相比仅使用真实数据训练,本框架性能提升超92%;引入AdaMix后进一步提升17%,验证了其在提升策略泛化方面的有效性。
原文摘要 · Abstract (English)
The generalization of vision-language-action (VLA) models heavily relies on diverse training data. However, acquiring large-scale data for robot manipulation across varied object appearances is costly and labor-intensive. To address this limitation, we introduce Embodied Manipulation Media Adaptation (EMMA), a framework for augmenting VLA policies that combines a generative data engine with an effective training pipeline. We introduce DreamTransfer, a diffusion Transformer-based architecture for generating multi-view consistent and geometrically grounded embodied manipulation videos. DreamTransfer enables visual editing of robot videos through prompts, allowing for changes to the foreground, background, and lighting while preserving their 3D structure and geometric validity. We also utilize a hybrid training set of real and generated data and propose AdaMix to enhance the training process. AdaMix is a training strategy that adaptively weights samples according to policy performance to emphasize challenging samples. Comprehensive evaluations demonstrate that videos created by DreamTransfer yield substantial improvements over previous video generation techniques in multi-view consistency, geometric accuracy, and text-conditioning precision. We conduct extensive evaluations with a total of more than 1800 trials in both simulated and real-world robotic environments. In real-world robotic tasks with zero-shot visual settings, our framework achieves a relative performance increase of over 92% compared to training with real data alone, and improves by an additional 17% with AdaMix, demonstrating its efficacy in enhancing policy generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。