arXiv:2502.15438cs.CV2025-02

解决自动驾驶视觉占位网络的闪烁问题,轻量级提升时序一致性

Deflickering Vision-Based Occupancy Networks through Lightweight Spatio-Temporal Correlation

  • 通过双交叉注意力融合历史静态与运动信息,生成修正占位成分
  • 在两个基准数据集上显著减少闪烁现象,计算开销极低
  • 适合追求高时序稳定性的自动驾驶感知系统使用

基于视觉的占位网络(VONs)为自动驾驶中三维环境重建提供了端到端解决方案。然而,现有方法常因时序不一致导致闪烁效应,损害时序连贯性并影响下游决策。尽管近期方法引入历史信息缓解该问题,但通常带来高计算成本,并可能引入错位或冗余特征干扰目标检测。本文提出OccLinker,一种可轻松集成至现有VONs的轻量级插件框架。该方法高效融合历史静态与运动线索,通过双交叉注意力机制学习当前特征与历史特征间的稀疏潜在关联,并生成修正占位成分以优化基础网络预测。此外,我们提出一种新的时序一致性度量指标,用于定量评估闪烁程度。在两个基准数据集上的大量实验表明,本方法在极低计算开销下实现更优性能,有效降低闪烁伪影。

原文摘要 · Abstract (English)

Vision-based occupancy networks (VONs) provide an end-to-end solution for reconstructing 3D environments in autonomous driving. However, existing methods often suffer from temporal inconsistencies, manifesting as flickering effects that degrade temporal coherence and adversely affect downstream decision-making. While recent approaches incorporate historical information to alleviate this issue, they often incur high computational costs and may introduce misaligned or redundant features that interfere with object detection. We propose OccLinker, a novel plugin framework that can be easily integrated into existing VONs to improve performance. Our method efficiently consolidates historical static and motion cues, learns sparse latent correlations with current features through a dual cross-attention mechanism, and generates correction occupancy components to refine the base network predictions. In addition, we introduce a new temporal consistency metric to quantitatively measure flickering effects. Extensive experiments on two benchmark datasets demonstrate that our method achieves superior performance with minimal computational overhead while effectively reducing flickering artifacts.

占位网络时序一致性自动驾驶轻量级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。