通过捕捉动态物体特征,提升基于视觉的3D语义场景补全精度。
Vision-based 3D Semantic Scene Completion via Capture Dynamic Representations
- 利用多模态模型对齐2D语义到3D空间,分离动态与静态特征。
- 在三个数据集上优于现有方法,动态物体干扰下仍保持高精度。
- 适合自动驾驶场景中复杂动态环境下的3D重建任务。
基于视觉的语义场景补全任务旨在从2D图像中预测密集的几何与语义3D场景表征。然而,场景中的动态物体严重影响模型从2D图像推断3D结构的准确性。现有方法简单堆叠多帧图像以增加语义信息,却忽略了动态物体和无纹理区域破坏多视角一致性与匹配可靠性的问题。为此,我们提出CDScene:一种基于视觉的鲁棒语义场景补全方法,通过捕捉动态表征实现改进。首先,利用多模态大规模模型提取2D显式语义并映射至3D空间;其次,结合单目与双目深度特性,将场景信息解耦为动态特征(包含动态物体周围结构关系)与静态特征(包含密集上下文空间信息);最后设计动态-静态自适应融合模块,有效提取并聚合互补特征,在自动驾驶场景中实现鲁棒且精确的语义场景补全。在SemanticKITTI、SSCBench-KITTI360和SemanticKITTI-C数据集上的大量实验表明,CDScene在性能与鲁棒性上均优于当前最优方法。
原文摘要 · Abstract (English)
The vision-based semantic scene completion task aims to predict dense geometric and semantic 3D scene representations from 2D images. However, the presence of dynamic objects in the scene seriously affects the accuracy of the model inferring 3D structures from 2D images. Existing methods simply stack multiple frames of image input to increase dense scene semantic information, but ignore the fact that dynamic objects and non-texture areas violate multi-view consistency and matching reliability. To address these issues, we propose a novel method, CDScene: Vision-based Robust Semantic Scene Completion via Capturing Dynamic Representations. First, we leverage a multimodal large-scale model to extract 2D explicit semantics and align them into 3D space. Second, we exploit the characteristics of monocular and stereo depth to decouple scene information into dynamic and static features. The dynamic features contain structural relationships around dynamic objects, and the static features contain dense contextual spatial information. Finally, we design a dynamic-static adaptive fusion module to effectively extract and aggregate complementary features, achieving robust and accurate semantic scene completion in autonomous driving scenarios. Extensive experimental results on the SemanticKITTI, SSCBench-KITTI360, and SemanticKITTI-C datasets demonstrate the superiority and robustness of CDScene over existing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。