用视觉互动意图指导3D物体功能定位,提升模型泛化能力
HAMMER: Harnessing MLLM via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
- 通过图像中的交互意图生成接触感知嵌入,驱动文本功能标签推断
- 在公开数据集和新构建的损坏基准上,3D功能定位准确率显著优于现有方法
- 适合研究多模态大模型、3D场景理解与具身智能的开发者
人类常通过图像或视频中观察到的交互行为识别3D物体的功能,一旦形成知识即可泛化到新物体。受此启发,我们提出HARMER框架,利用新兴的多模态大语言模型(MLLMs)实现基于交互意图的3D功能定位。不同于生成显式属性描述或依赖现成的2D分割器,我们聚合图像中呈现的交互意图,生成接触感知嵌入,并引导模型推断文本功能标签,确保充分挖掘物体语义与上下文线索。进一步设计分层跨模态融合机制,充分利用MLLM提供的互补信息以优化3D表示;引入多粒度几何提升模块,将空间特征注入提取的意图嵌入,从而实现精确的3D功能定位。在公开数据集及新构建的损坏基准上的大量实验表明,HARMER在性能和鲁棒性上均优于现有方法。所有代码与权重均已开源。
原文摘要 · Abstract (English)
Humans commonly identify 3D object affordance through observed interactions in images or videos, and once formed, such knowledge can be generically generalized to novel objects. Inspired by this principle, we advocate for a novel framework that leverages emerging multimodal large language models (MLLMs) for interaction intention-driven 3D affordance grounding, namely HAMMER. Instead of generating explicit object attribute descriptions or relying on off-the-shelf 2D segmenters, we alternatively aggregate the interaction intention depicted in the image into a contact-aware embedding and guide the model to infer textual affordance labels, ensuring it thoroughly excavates object semantics and contextual cues. We further devise a hierarchical cross-modal integration mechanism to fully exploit the complementary information from the MLLM for 3D representation refinement and introduce a multi-granular geometry lifting module that infuses spatial characteristics into the extracted intention embedding, thus facilitating accurate 3D affordance localization. Extensive experiments on public datasets and our newly constructed corrupted benchmark demonstrate the superiority and robustness of HAMMER compared to existing approaches. All code and weights are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。