用扩散模型先验学习3D物体操作属性,实现开放世界泛化。
Diffusion Models are Open-World Affordance Learners: Leveraging Generative Priors for 3D Affordance Learning
- 从文本到图像扩散模型中提取操作属性先验知识
- 在单样本设置下仍显著优于现有方法
- 适合需要开放世界泛化的机器人交互研究者
3D操作属性定位旨在理解多样化物体如何被操控,是具身交互的核心。然而,以往方法难以泛化至分布外的开放世界场景,导致数据集表现与真实应用需求之间存在巨大差距。受‘我无法创造,便不理解’启发,我们发现生成模型可生成语义合理的物-人交互图像,表明其内在编码了操作属性概念。基于此,我们提出DAG——首个基于扩散模型的3D操作属性定位框架,利用文本到图像扩散模型提取通用操作属性先验,用于3D操作属性预测。具体而言,我们从扩散模型中提取操作属性先验以编码物-人交互先验,并设计多源操作属性解码器实现密集3D操作属性预测。大量实验表明,DAG持续优于当前最优方法,在极具挑战性的单样本设置下也表现出强大开放世界泛化能力。代码已公开于https://github.com/hq-King/DAG。
原文摘要 · Abstract (English)
3D affordance grounding aims to understand how diverse objects can be manipulated, making it a cornerstone of embodied interaction. However, prior works struggle to generalize to out-of-distribution, open-world scenarios, leaving a critical gap between limited dataset performance and real-world application needs. Inspired by the saying: \textit{\textbf{``What I can not create, I do not understand''}}, we find generative models can generate semantically valid HOI images, which indicates inherent encoding of affordance concepts. Building on this insight, we propose DAG, the first innovative diffusion-based 3D affordance grounding framework that extracts general affordance knowledge from text-to-image diffusion models for 3D affordance prediction. Specifically, we extract the affordance priors from a diffusion model to encode HOI priors, and design an affordance block with a multi-source affordance decoder for dense 3D affordance prediction. Extensive experiments show that DAG consistently outperforms state-of-the-art methods and exhibits strong open-world generalization, even in the challenging one-shot setting. The code of our method is released on \textcolor{blue}{\textit{https://github.com/hq-King/DAG}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。