提出分层扩散模型A0,让机器人理解物体操作空间位置与方式。
A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
- 分两层建模:先理解操作空间,再执行具体动作。
- 在100万接触点数据上预训练,支持多平台通用操作。
- 适合复杂任务如擦板、堆叠,对机器人平台不敏感。
机器人操作面临理解空间可操作性的挑战——即物体交互的‘何处’与‘如何’,这对擦板、堆叠等复杂任务至关重要。现有模块化与端到端方法普遍缺乏稳健的空间推理能力。不同于基于点或流的密集空间表示或轨迹建模方法,本文提出A0:一种分层的可操作性感知扩散模型,将操作任务分解为高层空间可操作性理解与底层动作执行。A0采用无具身可操作性表征,通过预测接触点与接触后轨迹,捕捉以物体为中心的空间可操作性。模型在100万条接触点数据上预训练,并在标注轨迹上微调,实现跨平台泛化。关键组件包括用于运动感知特征提取的位置偏移注意力(Position Offset Attention)和用于精确坐标映射的空间信息聚合层(Spatial Information Aggregation Layer)。模型输出由动作执行模块解析执行。在Franka、Kinova、Realman和Dobot等多种机器人系统上的实验表明,A0在复杂任务中表现优越,展现出高效性、灵活性与真实场景适用性。
原文摘要 · Abstract (English)
Robotic manipulation faces critical challenges in understanding spatial affordances--the "where" and "how" of object interactions--essential for complex manipulation tasks like wiping a board or stacking objects. Existing methods, including modular-based and end-to-end approaches, often lack robust spatial reasoning capabilities. Unlike recent point-based and flow-based affordance methods that focus on dense spatial representations or trajectory modeling, we propose A0, a hierarchical affordance-aware diffusion model that decomposes manipulation tasks into high-level spatial affordance understanding and low-level action execution. A0 leverages the Embodiment-Agnostic Affordance Representation, which captures object-centric spatial affordances by predicting contact points and post-contact trajectories. A0 is pre-trained on 1 million contact points data and fine-tuned on annotated trajectories, enabling generalization across platforms. Key components include Position Offset Attention for motion-aware feature extraction and a Spatial Information Aggregation Layer for precise coordinate mapping. The model's output is executed by the action execution module. Experiments on multiple robotic systems (Franka, Kinova, Realman, and Dobot) demonstrate A0's superior performance in complex tasks, showcasing its efficiency, flexibility, and real-world applicability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。