用视觉语言模型提升相机3D语义场景补全的上下文理解能力
VLScene: Vision-Language Guidance Distillation for Camera-Based 3D Semantic Scene Completion
- 利用视觉语言模型注入高层语义先验,增强图像特征
- 在SemanticKITTI和SSCBench-KITTI-360上分别达17.52和19.10 mIoU
- 适合关注3D感知中语义与几何协同建模的研究者
基于摄像头的3D语义场景补全(SSC)为自动驾驶提供密集的几何与语义感知。然而,图像信息有限导致模型易受遮挡和透视畸变带来的几何模糊影响。现有方法往往缺乏对象间的显式语义建模,限制了对3D语义上下文的理解。为此,我们提出新方法VLScene:视觉语言引导蒸馏用于相机基3D语义场景补全。核心思想是利用视觉语言模型引入高层语义先验,提供3D场景理解所需的物体空间上下文。具体地,设计视觉语言引导蒸馏过程以增强图像特征,有效捕捉周围环境的语义知识,提升空间上下文推理能力。此外,引入几何-语义稀疏感知机制,传播邻域几何结构,并通过上下文稀疏交互增强语义信息。实验表明,VLScene在挑战性基准SemanticKITTI和SSCBench-KITTI-360上取得排名第一性能,mIoU分别达到17.52和19.10。
原文摘要 · Abstract (English)
Camera-based 3D semantic scene completion (SSC) provides dense geometric and semantic perception for autonomous driving. However, images provide limited information making the model susceptible to geometric ambiguity caused by occlusion and perspective distortion. Existing methods often lack explicit semantic modeling between objects, limiting their perception of 3D semantic context. To address these challenges, we propose a novel method VLScene: Vision-Language Guidance Distillation for Camera-based 3D Semantic Scene Completion. The key insight is to use the vision-language model to introduce high-level semantic priors to provide the object spatial context required for 3D scene understanding. Specifically, we design a vision-language guidance distillation process to enhance image features, which can effectively capture semantic knowledge from the surrounding environment and improve spatial context reasoning. In addition, we introduce a geometric-semantic sparse awareness mechanism to propagate geometric structures in the neighborhood and enhance semantic information through contextual sparse interactions. Experimental results demonstrate that VLScene achieves rank-1st performance on challenging benchmarks--SemanticKITTI and SSCBench-KITTI-360, yielding remarkably mIoU scores of 17.52 and 19.10, respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。