arXiv:2603.17647cs.CV2026-03

让AI理解3D物体功能区域,支持开放词汇下的精准定位。

Part-Aware Open-Vocabulary 3D Affordance Grounding via Prototypical Semantic and Geometric Alignment

  • 用大模型生成部件级指令补全语义,提升跨对象理解。
  • 通过原型聚合与关系建模,实现几何与语义的精细对齐。
  • 适用于机器人交互、智能助手等需要理解物体功能的场景。

将自然语言问题定位到3D物体中具有功能性的区域——即语言驱动的3D功能定位——是具身智能与人机交互的关键。现有方法虽从标注驱动转向语言驱动,仍面临开放词汇泛化能力弱、细粒度几何对齐差、部件级语义一致性不足等问题。为此,我们提出一种两阶段跨模态框架,增强语义与几何表征以实现开放词汇下的3D功能定位。第一阶段利用大语言模型生成部件感知指令,恢复缺失语义,使模型能关联语义相似的功能;第二阶段引入两个关键组件:功能原型聚合(APA),捕捉不同物体间同一功能的几何一致性;以及对象内关系建模(IORM),细化对象内部几何差异,支持精确语义对齐。我们在新提出的基准及两个现有基准上进行了广泛实验,结果表明该方法优于现有方法。

原文摘要 · Abstract (English)

Grounding natural language questions to functionally relevant regions in 3D objects -- termed language-driven 3D affordance grounding -- is essential for embodied intelligence and human-AI interaction. Existing methods, while progressing from label-based to language-driven approaches, still face challenges in open-vocabulary generalization, fine-grained geometric alignment, and part-level semantic consistency. To address these issues, we propose a novel two-stage cross-modal framework that enhances both semantic and geometric representations for open-vocabulary 3D affordance grounding. In the first stage, large language models generate part-aware instructions to recover missing semantics, enabling the model to link semantically similar affordances. In the second stage, we introduce two key components: Affordance Prototype Aggregation (APA), which captures cross-object geometric consistency for each affordance, and Intra-Object Relational Modeling (IORM), which refines geometric differentiation within objects to support precise semantic alignment. We validate the effectiveness of our method through extensive experiments on a newly introduced benchmark, as well as two existing benchmarks, demonstrating superior performance in comparison with existing methods.

3D理解语言对齐具身智能开放词汇

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。