用光流引导时序信息融合,提升3D语义场景补全精度
Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance
- 通过光流对齐多帧特征,捕捉运动感知上下文
- 在SemanticKITTI和SSCBench上达到当前最优性能
- 适合自动驾驶场景理解与时序感知的视觉任务
3D语义场景补全(SSC)为自动驾驶感知提供完整的场景几何与语义信息,对实现准确可靠的决策至关重要。然而现有方法仅依赖单帧稀疏信息或简单堆叠多帧特征,难以获取有效场景上下文,忽略关键运动动态,难以保证时序一致性。为此,本文提出FlowScene:基于光流引导的时序3D语义场景补全方法。该方法利用光流融合运动、多视角、遮挡等上下文信息,显著提升3D场景补全精度。具体包含两个核心组件:(1) 光流引导的时序聚合模块,通过光流对齐并聚合时序特征,捕捉运动感知上下文与可变形结构;(2) 遮挡引导的体素精炼模块,将遮挡掩码与时序聚合特征注入3D体素空间,自适应优化体素表示以实现显式几何建模。实验表明,FlowScene在SemanticKITTI和SSCBench-KITTI-360基准上均取得领先性能。
原文摘要 · Abstract (English)
3D Semantic Scene Completion (SSC) provides comprehensive scene geometry and semantics for autonomous driving perception, which is crucial for enabling accurate and reliable decision-making. However, existing SSC methods are limited to capturing sparse information from the current frame or naively stacking multi-frame temporal features, thereby failing to acquire effective scene context. These approaches ignore critical motion dynamics and struggle to achieve temporal consistency. To address the above challenges, we propose a novel temporal SSC method FlowScene: Learning Temporal 3D Semantic Scene Completion via Optical Flow Guidance. By leveraging optical flow, FlowScene can integrate motion, different viewpoints, occlusions, and other contextual cues, thereby significantly improving the accuracy of 3D scene completion. Specifically, our framework introduces two key components: (1) a Flow-Guided Temporal Aggregation module that aligns and aggregates temporal features using optical flow, capturing motion-aware context and deformable structures; and (2) an Occlusion-Guided Voxel Refinement module that injects occlusion masks and temporally aggregated features into 3D voxel space, adaptively refining voxel representations for explicit geometric modeling. Experimental results demonstrate that FlowScene achieves state-of-the-art performance on the SemanticKITTI and SSCBench-KITTI-360 benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。