arXiv:2410.11564cs.ROcs.CV2024-10被引 7

用视觉语言模型提升点云物体功能理解,让机器人更懂怎么操作物体。

PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model

  • 融合点云几何与大模型语义,用提示工程增强文本上下文。
  • 在3D-AffordanceNet上对完整和部分点云均超越基线方法。
  • 适合研究机器人交互、3D场景理解与多模态学习的开发者。

功能理解旨在识别3D物体上可操作的区域,对机器人与物理世界交互至关重要。尽管视觉语言模型在高层次推理与长程规划方面表现优异,但在掌握有效人机交互所需的细微物理属性方面仍存在不足。本文提出PAVLM(点云功能视觉语言模型),利用预训练语言模型中丰富的多模态知识,增强点云的3D功能理解。PAVLM通过几何引导传播模块与大语言模型(LLM)的隐含嵌入融合,丰富视觉语义。在语言端,我们使用Llama-3.1模型生成上下文感知的细化文本,为指令输入注入深层语义线索。在3D-AffordanceNet基准测试中,PAVLM在完整与部分点云上均优于基线方法,尤其在新开放世界功能任务的泛化能力突出。

原文摘要 · Abstract (English)

Affordance understanding, the task of identifying actionable regions on 3D objects, plays a vital role in allowing robotic systems to engage with and operate within the physical world. Although Visual Language Models (VLMs) have excelled in high-level reasoning and long-horizon planning for robotic manipulation, they still fall short in grasping the nuanced physical properties required for effective human-robot interaction. In this paper, we introduce PAVLM (Point cloud Affordance Vision-Language Model), an innovative framework that utilizes the extensive multimodal knowledge embedded in pre-trained language models to enhance 3D affordance understanding of point cloud. PAVLM integrates a geometric-guided propagation module with hidden embeddings from large language models (LLMs) to enrich visual semantics. On the language side, we prompt Llama-3.1 models to generate refined context-aware text, augmenting the instructional input with deeper semantic cues. Experimental results on the 3D-AffordanceNet benchmark demonstrate that PAVLM outperforms baseline methods for both full and partial point clouds, particularly excelling in its generalization to novel open-world affordance tasks of 3D objects. For more information, visit our project site: pavlm-source.github.io.

3D理解视觉语言模型机器人交互点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。