用生成式重建补全物体3D形状,让交互区域定位更完整。
Affostruction: 3D Affordance Grounding with Generative Reconstruction
- 通过多视角特征融合实现恒定复杂度的3D生成重建
- 在基准测试中达19.1 aIoU,3D重建达32.67 IoU
- 适合需要完整交互区域理解的机器人抓取任务
本文解决从物体的RGBD图像中进行交互属性定位的问题,目标是找出对应文本描述的动作所涉及的表面区域。现有方法仅在可见表面上预测交互区域,而本文提出Affostruction,一种生成式框架,能从部分观测的RGBD数据中重建完整的物体几何,并在完整形状上定位包括未观测区域在内的交互属性。该方法引入了多视角特征的稀疏体素融合以实现恒定复杂度的生成重建,采用基于流的建模捕捉交互分布的固有模糊性,并设计了由预测交互引导的主动视角选择策略。在多个挑战性基准上,Affostruction显著优于现有方法,交互属性定位达到19.1 aIoU,3D重建达到32.67 IoU。
原文摘要 · Abstract (English)
This paper addresses the problem of affordance grounding from RGBD images of an object, which aims to localize surface regions corresponding to a text query that describes an action on the object. While existing methods predict affordance regions only on visible surfaces, we propose Affostruction, a generative framework that reconstructs complete object geometry from partial RGBD observations and grounds affordances on the full shape including unobserved regions. Our approach introduces sparse voxel fusion of multi-view features for constant-complexity generative reconstruction, a flow-based formulation that captures the inherent ambiguity of affordance distributions, and an active view selection strategy guided by predicted affordances. Affostruction outperforms existing methods by large margins on challenging benchmarks, achieving 19.1 aIoU on affordance grounding and 32.67 IoU for 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。