arXiv:2501.04995cs.CVcs.AI2025-01AAAI被引 9

用图像增强3D指代分割,解决信息丢失和指令模糊问题

IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation

  • 融合多视角图像补偿点云信息损失
  • 通过表达与视觉交互生成任务导向信号
  • 在两个基准上分别提升1.9和4.2点

3D指代表达分割(3D-RES)旨在根据给定表达分割点云场景。现有方法面临两大挑战:特征模糊性与意图模糊性。特征模糊性源于点云获取中因光照、视角等限制导致的信息损失或失真;意图模糊性则指模型在解码过程中对所有查询一视同仁,缺乏自上而下的任务特定引导。本文提出图像增强提示解码网络(IPDN),利用多视角图像与任务驱动信息提升模型推理能力。为解决特征模糊性,提出多视角语义嵌入(MSE)模块,将多视角2D图像信息注入3D场景以弥补空间信息损失。为应对意图模糊性,设计提示感知解码器(PAD),通过表达与视觉特征的交互提取任务驱动信号,引导解码过程。大量实验表明,IPDN在3D-RES与3D-GRES任务上的mIoU指标分别优于当前最优方法1.9和4.2个百分点。

原文摘要 · Abstract (English)

3D Referring Expression Segmentation (3D-RES) aims to segment point cloud scenes based on a given expression. However, existing 3D-RES approaches face two major challenges: feature ambiguity and intent ambiguity. Feature ambiguity arises from information loss or distortion during point cloud acquisition due to limitations such as lighting and viewpoint. Intent ambiguity refers to the model's equal treatment of all queries during the decoding process, lacking top-down task-specific guidance. In this paper, we introduce an Image enhanced Prompt Decoding Network (IPDN), which leverages multi-view images and task-driven information to enhance the model's reasoning capabilities. To address feature ambiguity, we propose the Multi-view Semantic Embedding (MSE) module, which injects multi-view 2D image information into the 3D scene and compensates for potential spatial information loss. To tackle intent ambiguity, we designed a Prompt-Aware Decoder (PAD) that guides the decoding process by deriving task-driven signals from the interaction between the expression and visual features. Comprehensive experiments demonstrate that IPDN outperforms the state-ofthe-art by 1.9 and 4.2 points in mIoU metrics on the 3D-RES and 3D-GRES tasks, respectively.

3D分割指代表达图像增强多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。