用可组合的扩散模型生成机器人任务数据,零样本泛化到新任务组合。
Iterative Compositional Data Generation for Robot Control
- 将任务分解为机器人、物体等组件,通过注意力学习交互关系。
- 仅用少量训练任务,就能零样本生成高质量动作序列,覆盖近90%未见任务。
- 通过强化学习迭代优化合成数据,让模型自进化,适合多任务机器人研究者。
机器人操作数据采集成本高,在多物体、多机器人、多环境场景下难以获取所有任务的示范数据。尽管现有生成模型能为单一任务合成有效数据,却无法利用任务的组合结构,且难以泛化至未见的任务组合。本文提出一种语义可组合的扩散变换器,将状态转移分解为机器人、物体、障碍物和目标特定组件,并通过注意力机制学习其交互。在有限任务子集上训练后,该模型可零样本生成高质量状态转移序列,用于学习未见任务组合的控制策略。进一步引入基于离线强化学习的迭代自优化流程,验证并回填合成数据以改进后续训练。相比传统单体模型与硬编码组合基线,本方法显著提升零样本性能,最终解决近全部保留任务,且学习表征中涌现出有意义的组合结构。
原文摘要 · Abstract (English)
Collecting robotic manipulation data is expensive, making it impractical to acquire demonstrations for the combinatorially large space of tasks that arise in multi-object, multi-robot, and multi-environment settings. While recent generative models can synthesize useful data for individual tasks, they do not exploit the compositional structure of robotic domains and struggle to generalize to unseen task combinations. We propose a semantic compositional diffusion transformer that factorizes transitions into robot-, object-, obstacle-, and objective-specific components and learns their interactions through attention. Once trained on a limited subset of tasks, we show that our model can zero-shot generate high-quality transitions from which we can learn control policies for unseen task combinations. Then, we introduce an iterative self-improvement procedure in which synthetic data is validated via offline reinforcement learning and incorporated into subsequent training rounds. Our approach substantially improves zero-shot performance over monolithic and hard-coded compositional baselines, ultimately solving nearly all held-out tasks and demonstrating the emergence of meaningful compositional structure in the learned representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。