arXiv:2506.19312cs.CVcs.AI2025-06

用语言模型提升3D点云与文本的精细对齐,增强物体功能识别能力。

Capturing Fine-Grained Alignments Improves 3D Affordance Detection

  • 引入预训练语言模型捕捉点云与文本的细粒度对齐
  • 在3D AffordanceNet上准确率和平均交并比均优于现有方法
  • 适合研究3D视觉与语言融合、智能机器人交互的开发者

本文针对3D点云中的物体功能检测难题,提出一种新方法LM-AD。该任务需精准建模点云与文本间的细粒度对齐,而现有方法多依赖点云与文本嵌入间的简单余弦相似度,表达力不足。为此,我们设计了可高效捕捉对齐关系的先验语言模型增强模块(Affordance Query Module, AQM)。实验表明,所提方法在3D AffordanceNet数据集上,于准确率与平均交并比(mIoU)指标上均超越现有方法。

原文摘要 · Abstract (English)

In this work, we address the challenge of affordance detection in 3D point clouds, a task that requires effectively capturing fine-grained alignments between point clouds and text. Existing methods often struggle to model such alignments, resulting in limited performance on standard benchmarks. A key limitation of these approaches is their reliance on simple cosine similarity between point cloud and text embeddings, which lacks the expressiveness needed for fine-grained reasoning. To address this limitation, we propose LM-AD, a novel method for affordance detection in 3D point clouds. Moreover, we introduce the Affordance Query Module (AQM), which efficiently captures fine-grained alignment between point clouds and text by leveraging a pretrained language model. We demonstrated that our method outperformed existing approaches in terms of accuracy and mean Intersection over Union on the 3D AffordanceNet dataset.

3D感知视觉语言对齐功能检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。