arXiv:2603.21944cs.CV2026-03

用语义约束融合多视角3D点云,避免误合并与碎片化。

Group3D: MLLM-Driven Semantic Grouping for Open-Vocabulary 3D Object Detection

  • 将多模态大模型的语义知识融入实例构建过程
  • 在ScanNet和ARKitScenes上达到最新性能,零样本泛化强
  • 适合关注开放词汇3D检测与跨视角一致性研究者

开放词汇3D目标检测旨在定位并识别超出固定训练类别的物体。在多视角RGB设置中,现有方法通常将几何实例构建与语义标注解耦,生成无类别片段并在后期分配开放词汇类别。这种解耦使实例构建主要依赖几何一致性,缺乏语义约束,导致视图依赖且不完整的几何证据引发不可逆的关联错误,如不同物体的过合并或单个实例的碎片化。我们提出Group3D,一个整合语义约束至实例构建过程的多视角开放词汇3D检测框架。Group3D利用多模态大语言模型(MLLM)生成场景自适应词汇,并将其组织为编码合理跨视图类别等价性的语义兼容组。这些组作为合并时的约束:仅当3D片段同时满足语义兼容性和几何一致性时才可关联。这种语义门控合并机制有效缓解了几何驱动的过合并问题,同时吸收多视角类别变异性。Group3D支持已知位姿与无位姿两种设置,仅依赖RGB观测。在ScanNet和ARKitScenes上的实验表明,Group3D在多视角开放词汇3D检测中达到最新性能,且在零样本场景下表现出强泛化能力。

原文摘要 · Abstract (English)

Open-vocabulary 3D object detection aims to localize and recognize objects beyond a fixed training taxonomy. In multi-view RGB settings, recent approaches often decouple geometry-based instance construction from semantic labeling, generating class-agnostic fragments and assigning open-vocabulary categories post hoc. While flexible, such decoupling leaves instance construction governed primarily by geometric consistency, without semantic constraints during merging. When geometric evidence is view-dependent and incomplete, this geometry-only merging can lead to irreversible association errors, including over-merging of distinct objects or fragmentation of a single instance. We propose Group3D, a multi-view open-vocabulary 3D detection framework that integrates semantic constraints directly into the instance construction process. Group3D maintains a scene-adaptive vocabulary derived from a multimodal large language model (MLLM) and organizes it into semantic compatibility groups that encode plausible cross-view category equivalence. These groups act as merge-time constraints: 3D fragments are associated only when they satisfy both semantic compatibility and geometric consistency. This semantically gated merging mitigates geometry-driven over-merging while absorbing multi-view category variability. Group3D supports both pose-known and pose-free settings, relying only on RGB observations. Experiments on ScanNet and ARKitScenes demonstrate that Group3D achieves state-of-the-art performance in multi-view open-vocabulary 3D detection, while exhibiting strong generalization in zero-shot scenarios. The project page is available at https://ubin108.github.io/Group3D/.

3D检测开放词汇多模态语义融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。