arXiv:2502.19247cs.CV2025-02CVPR被引 6

用代理注意力优化点云结构,提升3D视觉定位精度与速度。

ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding

  • 通过可变形聚类识别目标区域子流形,再用多模态代理引导变换。
  • 在易目标上提升7.49%、难目标上提升4.60%,计算量减少40.6%。
  • 适合实时3D视觉定位任务,尤其对噪声和冗余数据鲁棒。

具身智能要求代理基于语言指令实时交互3D环境,核心任务为以我为中心的3D视觉定位。然而,从RGB-D图像生成的点云包含大量冗余背景数据和固有噪声,干扰目标区域的流形结构。现有增强方法通常流程繁琐,不适用于实时任务。本文提出适用于多模态任务的代理变换(Proxy Transformation),高效优化点云流形。首先使用可变形点聚类识别目标区域的子流形;随后设计代理注意力模块,利用多模态代理指导点云变换。基于此,构建子流形变换生成模块:文本信息全局引导不同子流形的平移向量,优化目标区域间相对空间关系;图像信息则指导每个子流形内的线性变换,细化局部流形。大量实验表明,该方法显著优于现有方法,在易目标上提升7.49%、难目标上提升4.60%,同时将注意力块计算开销降低40.6%。结果确立了当前最先进水平,验证了方法的有效性与鲁棒性。

原文摘要 · Abstract (English)

Embodied intelligence requires agents to interact with 3D environments in real time based on language instructions. A foundational task in this domain is ego-centric 3D visual grounding. However, the point clouds rendered from RGB-D images retain a large amount of redundant background data and inherent noise, both of which can interfere with the manifold structure of the target regions. Existing point cloud enhancement methods often require a tedious process to improve the manifold, which is not suitable for real-time tasks. We propose Proxy Transformation suitable for multimodal task to efficiently improve the point cloud manifold. Our method first leverages Deformable Point Clustering to identify the point cloud sub-manifolds in target regions. Then, we propose a Proxy Attention module that utilizes multimodal proxies to guide point cloud transformation. Built upon Proxy Attention, we design a submanifold transformation generation module where textual information globally guides translation vectors for different submanifolds, optimizing relative spatial relationships of target regions. Simultaneously, image information guides linear transformations within each submanifold, refining the local point cloud manifold of target regions. Extensive experiments demonstrate that Proxy Transformation significantly outperforms all existing methods, achieving an impressive improvement of 7.49% on easy targets and 4.60% on hard targets, while reducing the computational overhead of attention blocks by 40.6%. These results establish a new SOTA in ego-centric 3D visual grounding, showcasing the effectiveness and robustness of our approach.

3D视觉定位点云处理多模态流形优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。