arXiv:2506.04837cs.CV2025-06

用大模型直接根据自然语言指令生成3D点云分割掩码

OpenMaskDINO3D : Reasoning 3D Segmentation via Large Language Model

  • 通过引入SEG标记和对象标识符,实现点云与文本的对齐
  • 在ScanNet数据集上达到高精度3D实例分割效果
  • 适合需要自然语言交互的3D场景理解任务

尽管感知系统近年来在2D推理分割方面取得显著进展,但仍依赖显式人类指令或预定义类别来识别目标对象。这些系统已能基于复杂隐含查询文本在二维上下文中进行推理与理解,生成准确的分割掩码。然而,3D推理分割仍缺乏相应框架。本文提出OpenMaskDINO3D,一种面向全面3D理解与分割的大语言模型。该模型处理点云数据与文本提示,生成实例分割掩码,在多种3D任务中表现优异。通过引入SEG token和对象标识符,实现高精度3D分割掩码生成,使模型可直接从自然语言指令输出精确的点云分割结果。在大规模ScanNet数据集上的实验验证了OpenMaskDINO3D在各类任务中的有效性。

原文摘要 · Abstract (English)

Although perception systems have made remarkable advancements in recent years, particularly in 2D reasoning segmentation, these systems still rely on explicit human instruction or pre-defined categories to identify target objects before executing visual recognition tasks. Such systems have matured significantly, demonstrating the ability to reason and comprehend implicit user intentions in two-dimensional contexts, producing accurate segmentation masks based on complex and implicit query text. However, a comparable framework and structure for 3D reasoning segmentation remain absent. This paper introduces OpenMaskDINO3D, a LLM designed for comprehensive 3D understanding and segmentation. OpenMaskDINO3D processes point cloud data and text prompts to produce instance segmentation masks, excelling in many 3D tasks. By introducing a SEG token and object identifier, we achieve high-precision 3D segmentation mask generation, enabling the model to directly produce accurate point cloud segmentation results from natural language instructions. Experimental results on large-scale ScanNet datasets validate the effectiveness of our OpenMaskDINO3D across various tasks.

3D分割大模型点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。