用光流对齐时序信息,分阶段融合深度数据,提升单目3D语义场景补全精度。
CurriFlow: Curriculum-Guided Depth Fusion with Optical Flow-Based Temporal Alignment for 3D Semantic Scene Completion
- 基于光流实现多帧特征对齐,增强时序一致性与动态物体理解。
- 训练中由稀疏激光雷达深度渐进过渡到密集立体深度,提升几何鲁棒性。
- 结合SAM语义先验,强化体素级语义学习,适合自动驾驶感知场景。
语义场景补全(SSC)旨在从单目图像中推断完整的3D几何与语义信息,是自动驾驶视觉感知的关键能力。然而,现有方法依赖时序堆叠或深度投影,缺乏显式的运动推理,难以处理遮挡和噪声深度监督。本文提出CurriFlow,一种融合光流引导时序对齐与课程学习式深度融合的语义占用预测框架。CurriFlow采用多层级融合策略,利用预训练光流对齐分割、视觉与深度特征,提升时序一致性和动态物体理解。为增强几何鲁棒性,引入课程学习机制,在训练中逐步从稀疏但精确的激光雷达深度过渡到密集但噪声较大的立体深度,确保稳定优化并适应真实部署。此外,来自分割任意模型(SAM)的语义先验提供类别无关监督,加强体素级语义学习与空间一致性。在SemanticKITTI基准上的实验表明,CurriFlow达到16.9的平均交并比,验证了其运动引导与课程感知设计的有效性。
原文摘要 · Abstract (English)
Semantic Scene Completion (SSC) aims to infer complete 3D geometry and semantics from monocular images, serving as a crucial capability for camera-based perception in autonomous driving. However, existing SSC methods relying on temporal stacking or depth projection often lack explicit motion reasoning and struggle with occlusions and noisy depth supervision. We propose CurriFlow, a novel semantic occupancy prediction framework that integrates optical flow-based temporal alignment with curriculum-guided depth fusion. CurriFlow employs a multi-level fusion strategy to align segmentation, visual, and depth features across frames using pre-trained optical flow, thereby improving temporal consistency and dynamic object understanding. To enhance geometric robustness, a curriculum learning mechanism progressively transitions from sparse yet accurate LiDAR depth to dense but noisy stereo depth during training, ensuring stable optimization and seamless adaptation to real-world deployment. Furthermore, semantic priors from the Segment Anything Model (SAM) provide category-agnostic supervision, strengthening voxel-level semantic learning and spatial consistency. Experiments on the SemanticKITTI benchmark demonstrate that CurriFlow achieves state-of-the-art performance with a mean IoU of 16.9, validating the effectiveness of our motion-guided and curriculum-aware design for camera-based 3D semantic scene completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。