arXiv:2602.22716cs.CVcs.AI2026-02被引 3

用球坐标改进3D视觉语言模型的位置编码,提升空间感知能力。

SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs

  • 将点云令牌映射到球坐标空间,统一建模位置与方向角。
  • 在多个3D场景基准上实现性能提升,尤其在方向感知任务中显著优于传统方法。
  • 适合需要精准空间理解的3D多模态应用,如机器人导航、自动驾驶。

基于大语言模型的3D大视觉语言模型(3D LVLMs)在多模态任务中取得了显著进展。然而,其继承的位置依赖建模机制——旋转位置编码(RoPE),在3D多模态理解中仍不理想。原始的RoPE无法保留点云数据的三维空间结构,且相对距离计算忽略了角度依赖性,限制了模型对视觉表征方向变化的捕捉能力。为此,本文提出球坐标位置编码(SoPE),将点云令牌索引映射至三维球坐标空间,统一建模空间位置与方向角。该方法保持了点云数据的固有几何结构,增强空间感知,生成更一致、更具表现力的几何表征。此外,引入多尺度频率混合策略,融合不同频率域的特征信息。在多个3D场景基准上的实验验证了该方法的有效性,真实场景部署实验进一步展示了其强大的泛化能力。

原文摘要 · Abstract (English)

3D Large Vision-Language Models (3D LVLMs) built upon Large Language Models (LLMs) have achieved remarkable progress across various multimodal tasks. However, their inherited position-dependent modeling mechanism, Rotary Position Embedding (RoPE), remains suboptimal for 3D multimodal understanding. The vanilla RoPE formulation fails to preserve essential three-dimensional spatial structures when encoding 3D tokens, and its relative distance computation overlooks angular dependencies, hindering the model's ability to capture directional variations in visual representations. To overcome these limitations, we introduce Spherical Coordinate-based Positional Embedding (SoPE). Our method maps point-cloud token indices into a 3D spherical coordinate space, enabling unified modeling of spatial locations and directional angles. This formulation preserves the inherent geometric structure of point-cloud data, enhances spatial awareness, and yields more consistent and expressive geometric representations for multimodal learning. In addition, we introduce a multi-scale frequency mixing strategy to fuse feature information across different frequency domains. Experimental results on multiple 3D scene benchmarks validate the effectiveness of our approach, while real-world deployment experiments further demonstrate its strong generalization capability.

3D视觉位置编码点云多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。