arXiv:2506.05689cs.CV2025-06被引 2

对比3D点云与视频特征在大模型中的表现,发现合理采样点可媲美视频表示。

Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models

  • 用点云特征增强视觉令牌,融合3D信息提升理解能力
  • 精心采样与排序的点结构性能媲美视频特征
  • 在多个基准上达到顶尖水平,结果复现性高

为多模态大语言模型(MLLMs)有效表示3D场景至关重要但具挑战性。现有方法通常仅依赖2D图像特征并采用不同标记化方式。本文系统研究3D标记结构,对比基于视频与基于点云的表示,同时保持模型主干和参数一致。提出新方法,通过Sonata预训练点变换器V3编码器引入3D点云特征以丰富视觉令牌。实验表明,融合显式3D特征显著提升性能;进一步证明,当点云被巧妙采样和排序时,点基标记结构可媲美视频基结构。两种结构的最佳模型在多个3D理解基准上取得当前最优结果。我们强调对标记结构的分析是核心贡献之一,并透明报告多次种子平均结果,认为此做法对领域稳健进步至关重要。

原文摘要 · Abstract (English)

Effectively representing 3D scenes for Multimodal Large Language Models (MLLMs) is crucial yet challenging. Existing approaches commonly only rely on 2D image features and use varied tokenization approaches. This work presents a rigorous study of 3D token structures, systematically comparing video-based and point-based representations while maintaining consistent model backbones and parameters. We propose a novel approach that enriches visual tokens by incorporating 3D point cloud features from a Sonata pretrained Point Transformer V3 encoder. Our experiments demonstrate that merging explicit 3D features significantly boosts performance. Furthermore, we show that point-based token structures can rival video-based ones when the points are cleverly sampled and ordered. Our best models from both structures achieve state-of-the-art results on multiple 3D understanding benchmarks. We emphasize our analysis of token structures as a key contribution, alongside transparent reporting of results averaged over multiple seeds, a practice we believe is vital for robust progress in the field.

3D理解点云大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。