arXiv:2502.04981cs.CV2025-02ICCV被引 15

用视觉语言模型自动标注3D场景语义占用,免人工标注。

AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting

论文配图:AutoOcc: Automatic Open-Ended Semantic Occupancy Annotation via Vision-Language Guided Gaussian Splatting
图 1 · 摘自论文原文
  • 通过视觉语言模型注意力图引导高斯点云重建
  • 无需人工标注,实现开放词表的语义占用生成
  • 支持静态与动态复杂场景,适用于自动驾驶等应用

从原始传感器数据中获取高质量3D语义占用仍是一项关键但具有挑战性的任务,通常需要大量人工标注。本文提出AutoOcc,一种以视觉为中心的自动化开放词表语义占用标注流水线,结合由视觉语言模型引导的可微高斯点云(differentiable Gaussian splatting)。我们将开放词表语义3D占用重建问题建模为:利用视觉语言模型和基础视觉模型的注意力图生成场景占用。设计语义感知高斯作为中间几何描述符,并提出累积高斯到体素点映射算法,实现高效且准确的占用标注。所提框架在无任何人工标签的情况下,优于现有自动化占用标注方法。AutoOcc还实现了开放词表语义占用的自动标注,在静态及动态复杂场景中均表现稳健。

原文摘要 · Abstract (English)

Obtaining high-quality 3D semantic occupancy from raw sensor data remains an essential yet challenging task, often requiring extensive manual labeling. In this work, we propose AutoOcc, a vision-centric automated pipeline for open-ended semantic occupancy annotation that integrates differentiable Gaussian splatting guided by vision-language models. We formulate the open-ended semantic 3D occupancy reconstruction task to automatically generate scene occupancy by combining attention maps from vision-language models and foundation vision models. We devise semantic-aware Gaussians as intermediate geometric descriptors and propose a cumulative Gaussian-to-voxel splatting algorithm that enables effective and efficient occupancy annotation. Our framework outperforms existing automated occupancy annotation methods without human labels. AutoOcc also enables open-ended semantic occupancy auto-labeling, achieving robust performance in both static and dynamically complex scenarios.

3D语义自动标注高斯点云视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。