arXiv:2507.13801cs.CVcs.AI2025-07被引 3

用预测未来帧扩展感知范围,提升单目语义场景补全效果

One Step Closer: Creating the Future to Boost Monocular Semantic Scene Completion

  • 通过预测伪未来帧拓展感知视野,融合时空信息增强几何一致性
  • 在SemanticKITTI和SSCBench-KITTI-360上达到当前最优性能
  • 适合需要处理遮挡与视野受限的自动驾驶场景研究者

近年来,视觉3D语义场景补全(SSC)因其能从单张2D图像推断完整的3D场景布局与语义,成为自动驾驶的关键感知任务。然而在真实交通场景中,大量场景被遮挡或超出相机视域——这是现有单目SSC方法难以解决的根本挑战。为此,我们提出创建未来语义场景补全(CF-SSC),一种新颖的时序SSC框架,利用伪未来帧预测来扩展模型的有效感知范围。该方法结合位姿与深度信息建立精确的3D对应关系,实现在3D空间中对过去、当前及预测未来帧的几何一致融合。不同于传统方法依赖简单特征堆叠,我们的3D感知架构通过显式建模时空关系,实现更鲁棒的场景补全。在SemanticKITTI与SSCBench-KITTI-360基准上的全面实验表明,本方法达到当前最优性能,验证了其在提升遮挡推理与3D场景补全精度方面的有效性。

原文摘要 · Abstract (English)

In recent years, visual 3D Semantic Scene Completion (SSC) has emerged as a critical perception task for autonomous driving due to its ability to infer complete 3D scene layouts and semantics from single 2D images. However, in real-world traffic scenarios, a significant portion of the scene remains occluded or outside the camera's field of view -- a fundamental challenge that existing monocular SSC methods fail to address adequately. To overcome these limitations, we propose Creating the Future SSC (CF-SSC), a novel temporal SSC framework that leverages pseudo-future frame prediction to expand the model's effective perceptual range. Our approach combines poses and depths to establish accurate 3D correspondences, enabling geometrically-consistent fusion of past, present, and predicted future frames in 3D space. Unlike conventional methods that rely on simple feature stacking, our 3D-aware architecture achieves more robust scene completion by explicitly modeling spatial-temporal relationships. Comprehensive experiments on SemanticKITTI and SSCBench-KITTI-360 benchmarks demonstrate state-of-the-art performance, validating the effectiveness of our approach, highlighting our method's ability to improve occlusion reasoning and 3D scene completion accuracy.

语义场景补全自动驾驶时空建模单目3D

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。