用实体中心的扩散模型,让机器人零样本泛化到新物体组合任务。
EC-Diffuser: Multi-Object Manipulation via Entity-Centric Behavior Generation
- 基于物体中心表征与实体级Transformer,从图像中学习动作
- 在未见物体组合上实现零样本泛化,支持更多物体数量
- 结合扩散模型捕捉多模态行为,适合复杂多物场景
物体操作是日常任务的常见组成部分,但从高维观测中学习操作仍具挑战,尤其在多物体环境中,状态空间和目标行为的组合复杂性显著增加。尽管近期方法利用大规模离线数据训练像素级模型并取得性能提升,但在受限网络与数据规模下,难以实现组合泛化。为此,我们提出一种新型行为克隆(BC)方法:采用物体中心表示与实体中心Transformer,结合基于扩散的优化,实现从离线图像数据的高效学习。该方法首先将观测分解为物体中心表示,再通过实体级Transformer在物体层面计算注意力,同时预测物体动态与智能体动作。结合扩散模型对多模态行为分布的建模能力,显著提升了多物体任务表现,并实现了零样本泛化至未见物体组合,包括训练中未出现的物体数量。视频演示详见:https://sites.google.com/view/ec-diffuser。
原文摘要 · Abstract (English)
Object manipulation is a common component of everyday tasks, but learning to manipulate objects from high-dimensional observations presents significant challenges. These challenges are heightened in multi-object environments due to the combinatorial complexity of the state space as well as of the desired behaviors. While recent approaches have utilized large-scale offline data to train models from pixel observations, achieving performance gains through scaling, these methods struggle with compositional generalization in unseen object configurations with constrained network and dataset sizes. To address these issues, we propose a novel behavioral cloning (BC) approach that leverages object-centric representations and an entity-centric Transformer with diffusion-based optimization, enabling efficient learning from offline image data. Our method first decomposes observations into an object-centric representation, which is then processed by our entity-centric Transformer that computes attention at the object level, simultaneously predicting object dynamics and the agent's actions. Combined with the ability of diffusion models to capture multi-modal behavior distributions, this results in substantial performance improvements in multi-object tasks and, more importantly, enables compositional generalization. We present BC agents capable of zero-shot generalization to tasks with novel compositions of objects and goals, including larger numbers of objects than seen during training. We provide video rollouts on our webpage: https://sites.google.com/view/ec-diffuser.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。