arXiv:2603.17753cs.CV2026-03AAAI

提出双层次注意力机制,提升复杂场景下3D指代定位与分割精度。

PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation

论文配图:PC-CrossDiff: Point-Cluster Dual-Level Cross-Modal Differential Attention for Unified 3D Referring and Segmentation
图 1 · 摘自论文原文
  • 通过点级与聚类级差分注意力,自适应挖掘隐式空间线索。
  • 在ScanRefer隐式子集上3D指代任务准确率提升10.16%。
  • 适合需要高精度3D视觉定位的复杂场景应用。

3D视觉定位旨在通过指代理解(3DREC)和分割(3DRES)两项核心任务,定位自然语言表达所指的物体。现有方法在简单单对象场景中表现良好,但在真实世界常见的复杂多对象场景中性能显著下降,限制了实际应用。主要挑战在于:难以解析对区分视觉相似物体至关重要的隐式定位线索,以及无法有效抑制共现物体带来的动态空间干扰。为此,本文提出PC-CrossDiff,一种统一处理3DREC与3DRES的双层级跨模态差分注意力框架。具体包括:(i) 点级差分注意力(PLDA)模块,通过文本与点云间的双向差分注意力,利用可学习权重自适应提取隐式定位线索,增强判别性表示;(ii) 聚类级差分注意力(CLDA)模块,建立层级注意力机制,通过定位感知的差分注意力块,自适应强化相关空间关系,抑制模糊或无关的空间关联。该方法在ScanRefer、NR3D和SR3D基准上达到最优性能。尤其在ScanRefer的隐式子集上,3DREC任务的[email protected]得分提升10.16%,凸显其解析隐式空间线索的强大能力。

原文摘要 · Abstract (English)

3D Visual Grounding (3DVG) aims to localize the referent of natural language referring expressions through two core tasks: Referring Expression Comprehension (3DREC) and Segmentation (3DRES). While existing methods achieve high accuracy in simple, single-object scenes, they suffer from severe performance degradation in complex, multi-object scenes that are common in real-world settings, hindering practical deployment. Existing methods face two key challenges in complex, multi-object scenes: inadequate parsing of implicit localization cues critical for disambiguating visually similar objects, and ineffective suppression of dynamic spatial interference from co-occurring objects, resulting in degraded grounding accuracy. To address these challenges, we propose PC-CrossDiff, a unified dual-task framework with a dual-level cross-modal differential attention architecture for 3DREC and 3DRES. Specifically, the framework introduces: (i) Point-Level Differential Attention (PLDA) modules that apply bidirectional differential attention between text and point clouds, adaptively extracting implicit localization cues via learnable weights to improve discriminative representation; (ii) Cluster-Level Differential Attention (CLDA) modules that establish a hierarchical attention mechanism to adaptively enhance localization-relevant spatial relationships while suppressing ambiguous or irrelevant spatial relations through a localization-aware differential attention block. Our method achieves state-of-the-art performance on the ScanRefer, NR3D, and SR3D benchmarks. Notably, on the Implicit subsets of ScanRefer, it improves the [email protected] score by +10.16% for the 3DREC task, highlighting its strong ability to parse implicit spatial cues.

3D视觉定位跨模态注意力指代理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。