arXiv:2607.00351cs.RO2026-07

通过合成新演示数据,让机器人模型学会组合已有技能完成更多动作。

Unleashing More Actions via Action Compositional Training for VLA Models

论文配图:Unleashing More Actions via Action Compositional Training for VLA Models
图 1 · 摘自论文原文
  • 用模型自身潜空间生成新的物理合理演示,自动扩充训练数据。
  • 在仿真任务中成功率显著提升,突破原数据分布限制。
  • 无需人工操作,适合想低成本扩展机器人泛化能力的研究者。

视觉-语言-动作(VLA)模型在机器人操作中表现优异,但标准训练方式常导致模型过度拟合特定行为模式,难以泛化到分布外场景,即使这些场景仅需已知子技能的全新组合。尽管扩大数据集可缓解过拟合,但高质量机器人数据采集成本高昂且耗时。为解决这一难题,我们提出ACT-VLA框架——一种离线数据增强方法,利用模型的隐式任务表示,直接从已有任务中合成新型、物理合理的演示用于策略训练。该方法无需额外人工数据采集,自动扩展训练分布并减轻过拟合。我们在模拟环境中评估该方法,结果表明,基线VLA模型因原始分布过拟合而泛化能力差,而使用合成数据训练的策略则显著提高成功率,验证了基于现有任务自动合成演示是实现VLA模型广泛泛化的有效、可扩展且高效路径。

原文摘要 · Abstract (English)

Vision-Language-Action models excel at robotic manipulation, driven by the scale and diversity of demonstration data. However, standard training paradigms often cause VLA models to severely overfit to specific behavioral patterns, rendering them unable to generalize to out-of-distribution scenarios even when those scenarios merely require novel combinations of identical sub-skills. While expanding datasets can mitigate this overfitting, acquiring high-quality robot data remains notoriously labor-intensive and cost-prohibitive. To resolve this impasse without expensive human teleoperation and to truly unleash more actions,i.e., enable VLA models to compose known sub-skills into a much broader set of executable behaviors beyond the original demonstrations-we propose ACT-VLA (Action Compositional Training for VLA Models), an offline data augmentation framework that leverages the model's latent task representations to synthesize novel, physically valid demonstrations directly from existing tasks for policy training. By eliminating additional manual data collection, our method automatically expands the training distribution and mitigates overfitting. We evaluate our approach on challenging manipulation tasks in simulation. Experiments demonstrate that while baseline VLA models generalize poorly due to original distribution overfitting, policies trained with our synthesized data achieve substantially higher success rates, validating that leveraging existing tasks for automated demonstration synthesis provides an effective, scalable, and data-efficient route to broadening VLA generalization.

机器人技能组合数据增强泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。