arXiv:2501.01163cs.CV2025-01CVPR被引 79

3D-LLaVA用点云直接对话,让模型更懂三维世界。

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

  • 用统一架构直接处理点云,无需复杂预处理
  • 在多个基准上表现优异,支持精细场景理解
  • 适合需要与3D环境交互的智能助手开发

当前3D大模型在基于三维视觉的对话和推理中展现出巨大潜力,但如何进一步提升其细粒度场景理解能力并支持灵活的人机交互仍是难题。本文提出3D-LLaVA,一种简洁而强大的3D大模型,旨在作为理解、推理和交互三维世界的智能助手。不同于依赖复杂流水线(如离线多视角特征提取或特定任务头)的现有方法,3D-LLaVA采用极简设计,仅以点云为输入。其核心是新型全向超点变压器(Omni Superpoint Transformer, OST),集成三项功能:(1)视觉特征选择器,将视觉标记转换并筛选;(2)视觉提示编码器,将交互式视觉提示嵌入视觉标记空间;(3)指代掩码解码器,根据文本描述生成3D掩码。该多功能OST通过混合预训练获得感知先验,并作为连接3D数据与大语言模型的桥梁。经过统一指令微调后,3D-LLaVA在多个基准测试中取得优异表现。

原文摘要 · Abstract (English)

Current 3D Large Multimodal Models (3D LMMs) have shown tremendous potential in 3D-vision-based dialogue and reasoning. However, how to further enhance 3D LMMs to achieve fine-grained scene understanding and facilitate flexible human-agent interaction remains a challenging problem. In this work, we introduce 3D-LLaVA, a simple yet highly powerful 3D LMM designed to act as an intelligent assistant in comprehending, reasoning, and interacting with the 3D world. Unlike existing top-performing methods that rely on complicated pipelines-such as offline multi-view feature extraction or additional task-specific heads-3D-LLaVA adopts a minimalist design with integrated architecture and only takes point clouds as input. At the core of 3D-LLaVA is a new Omni Superpoint Transformer (OST), which integrates three functionalities: (1) a visual feature selector that converts and selects visual tokens, (2) a visual prompt encoder that embeds interactive visual prompts into the visual token space, and (3) a referring mask decoder that produces 3D masks based on text description. This versatile OST is empowered by the hybrid pretraining to obtain perception priors and leveraged as the visual connector that bridges the 3D data to the LLM. After performing unified instruction tuning, our 3D-LLaVA reports impressive results on various benchmarks.

3D视觉多模态模型点云处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。