让3D场景理解更好懂:用语言模型提升透明/反光物体的语义渲染效果
3D Vision-Language Gaussian Splatting
- 分治处理视觉与语言模态,用融合机制增强语义表达
- 在开放词汇分割任务中超越现有方法,显著提升透明和反射物体识别准确率
- 适合做机器人感知、自动驾驶等需理解复杂场景的项目
近期3D重建技术和视觉-语言模型的发展推动了多模态3D场景理解的进步,广泛应用于机器人、自动驾驶及虚拟/增强现实。然而,现有方法简单将语义信息嵌入3D重建过程,未平衡视觉与语言模态,导致透明或反射物体的语义渲染不佳,并在颜色模态上出现过拟合。为此,我们提出一种3D视觉-语言高斯点阵模型,重点强化语言模态的表征学习。设计了一种新型跨模态光栅化器,通过模态融合与平滑语义指示符提升语义渲染质量;同时采用相机视角混合技术,增强已有视图与合成视图间的语义一致性,有效缓解过拟合问题。大量实验表明,该方法在开放词汇语义分割任务中达到当前最优性能,显著优于现有方法。
原文摘要 · Abstract (English)
Recent advancements in 3D reconstruction methods and vision-language models have propelled the development of multi-modal 3D scene understanding, which has vital applications in robotics, autonomous driving, and virtual/augmented reality. However, current multi-modal scene understanding approaches have naively embedded semantic representations into 3D reconstruction methods without striking a balance between visual and language modalities, which leads to unsatisfying semantic rasterization of translucent or reflective objects, as well as over-fitting on color modality. To alleviate these limitations, we propose a solution that adequately handles the distinct visual and semantic modalities, i.e., a 3D vision-language Gaussian splatting model for scene understanding, to put emphasis on the representation learning of language modality. We propose a novel cross-modal rasterizer, using modality fusion along with a smoothed semantic indicator for enhancing semantic rasterization. We also employ a camera-view blending technique to improve semantic consistency between existing and synthesized views, thereby effectively mitigating over-fitting. Extensive experiments demonstrate that our method achieves state-of-the-art performance in open-vocabulary semantic segmentation, surpassing existing methods by a significant margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。