提出可泛化折叠衣物的视觉语言动作模型,提升机器人处理变形物体能力。
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation

- 基于视觉语言模型与动作专家,用流匹配生成连续动作。
- 在5000小时真实双臂示范数据上预训练,实现跨类别折叠泛化。
- 通过人机协同数据聚合提升纠错能力,适合家庭服务机器人使用。
现实世界中的家用机器人需要具备视觉-语言-动作(VLA)基础模型,以在多样物体、任务条件和家居环境中复用操作技能。衣物折叠是典型挑战,需应对不同类别、几何形状、材质和场景下的随机初始状态。现有VLA系统通常为不同物体类别训练独立策略,而混合多任务训练常因任务干扰导致性能下降。为此,本文提出DeMaVLA,一种面向可泛化变形物体操作的VLA基础模型。DeMaVLA采用视觉语言模型(VLM)主干网络并集成动作专家,通过流匹配方法实现连续动作生成。为提高效率,动作专家通过剪枝每层交替的Transformer层构建,同时保持与VLM主干的逐层对齐,显著降低训练与推理开销。DeMaVLA首先在约5000小时精选的真实世界双臂演示数据上进行预训练,获取通用操作先验;随后在融合自收集演示与真人介入纠正轨迹的混合折叠数据上进行后训练,通过人机协同数据聚合(DAgger)管道完成。实验表明,DeMaVLA在RoboTwin 2.0上表现优异,在自建家居折叠基准上实现强真实世界性能。结果凸显了可扩展真实数据、高效动作生成与纠正学习对通用型VLA策略在变形物体操作中的价值。
原文摘要 · Abstract (English)
Real-world household robots require Vision-Language-Action (VLA) foundation models that can acquire reusable manipulation skills across diverse objects, task conditions, and household environments. Deformable-object folding is a representative challenge, requiring robots to handle clothing items from random initial states across varying categories, geometries, materials, and scenes. However, existing VLA systems commonly train separate policies for different object categories, while naively mixed multi-task training often suffers from task interference and degraded performance. To move beyond category-specific folding policies, we introduce DeMaVLA, a VLA foundation model for generalizable Deformable Manipulation. DeMaVLA adopts a VLM backbone with an action expert and formulates continuous action generation using flow matching. To improve efficiency, the action expert is constructed by pruning every other transformer layer while preserving layer-wise alignment with the VLM backbone, reducing training and inference cost. DeMaVLA is first pre-trained on approximately 5,000 hours of selected real-world dual-arm demonstrations to acquire general manipulation priors. It is then post-trained on mixed folding data that aggregates self-collected demonstrations and corrective trajectories from real-robot failures across multiple folding tasks through a human-in-the-loop Data Aggregation~(DAgger) pipeline. Experiments show that DeMaVLA achieves competitive performance on RoboTwin 2.0 and strong real-world results on our household folding benchmark. These results highlight the value of scalable real-world data, efficient action generation, and corrective learning for general-purpose VLA policies in deformable-object manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。