arXiv:2507.02250cs.CV2025-07

用流匹配与选择性状态空间模型提升少帧3D占位预测精度

FMOcc: TPV-Driven Flow Matching for 3D Occupancy Prediction with Selective State Space Model

  • 基于流匹配的特征修复模块,补全缺失视觉信息
  • 在Occ3D-nuScenes上达43.1% RayIoU、39.8% mIoU,仅需两帧输入
  • 适合自动驾驶中远距离遮挡场景的高效占位预测

3D语义占位预测在自动驾驶中至关重要。然而,少帧图像固有局限与3D空间冗余会降低遮挡及远距离场景的预测精度。现有方法通过融合历史帧数据提升性能,但需额外数据与大量计算资源。本文提出FMOcc,一种基于三视角图(TPV)的占位预测网络,结合流匹配选择性状态空间模型。首先设计基于流匹配的特征修复模块(FMSSM),生成缺失特征;其次引入TPV SSM层与平面选择性状态空间模型(PS3M),有选择地过滤TPV特征,减少空格体对非空格体的影响,提升远距离场景效率与预测能力;最后采用掩码训练(MT)增强鲁棒性并应对传感器数据丢失问题。在Occ3D-nuScenes和OpenOcc数据集上的实验表明,FMOcc显著优于现有方法。使用两帧输入时,在Occ3D-nuScenes验证集上达到43.1% RayIoU与39.8% mIoU,OpenOcc上达42.6% RayIoU,推理内存仅5.4G,耗时330ms。

原文摘要 · Abstract (English)

3D semantic occupancy prediction plays a pivotal role in autonomous driving. However, inherent limitations of fewframe images and redundancy in 3D space compromise prediction accuracy for occluded and distant scenes. Existing methods enhance performance by fusing historical frame data, which need additional data and significant computational resources. To address these issues, this paper propose FMOcc, a Tri-perspective View (TPV) refinement occupancy network with flow matching selective state space model for few-frame 3D occupancy prediction. Firstly, to generate missing features, we designed a feature refinement module based on a flow matching model, which is called Flow Matching SSM module (FMSSM). Furthermore, by designing the TPV SSM layer and Plane Selective SSM (PS3M), we selectively filter TPV features to reduce the impact of air voxels on non-air voxels, thereby enhancing the overall efficiency of the model and prediction capability for distant scenes. Finally, we design the Mask Training (MT) method to enhance the robustness of FMOcc and address the issue of sensor data loss. Experimental results on the Occ3D-nuScenes and OpenOcc datasets show that our FMOcc outperforms existing state-of-theart methods. Our FMOcc with two frame input achieves notable scores of 43.1% RayIoU and 39.8% mIoU on Occ3D-nuScenes validation, 42.6% RayIoU on OpenOcc with 5.4 G inference memory and 330ms inference time.

3D占位流匹配自动驾驶状态空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。