arXiv:2503.06992cs.CV2025-03CVPR被引 6

用共同时空梯度桥接帧与事件数据,提升高速场景光流精度。

Bridge Frame and Event: Common Spatiotemporal Fusion for High-Dynamic Scene Optical Flow

  • 构建共用时空梯度作为桥梁,对齐帧与事件的特征分布。
  • 在MVS-Benchmark上达到4.39的EPE,优于现有方法。
  • 适合需要高动态光流的自动驾驶与机器人视觉任务。

高速场景光流估计面临因大位移导致的空间模糊和时间不连续运动问题,恶化了光流的时空特征。现有方法多直接融合帧与事件数据,但因模态间异构表征差异大,效果有限。为此,本文提出一种新型共用时空融合框架,包含视觉边界定位与运动相关性融合。具体地,在视觉边界定位中,发现帧与事件共享相似的时空梯度,其分布与提取的边界一致,据此设计共用时空梯度以约束边界定位。在运动相关性融合中,观察到基于帧的运动具有空间密集但时间不连续的相关性,而基于事件的运动则为稀疏但时间连续,因此利用参考边界引导两种模态的互补运动知识融合。该方法不仅能缓解跨模态特征差异,还能使融合过程可解释,生成稠密且连续的光流。大量实验验证了其优越性,尤其在MVS-Benchmark上取得4.39的EPE(End-point Error)。

原文摘要 · Abstract (English)

High-dynamic scene optical flow is a challenging task, which suffers spatial blur and temporal discontinuous motion due to large displacement in frame imaging, thus deteriorating the spatiotemporal feature of optical flow. Typically, existing methods mainly introduce event camera to directly fuse the spatiotemporal features between the two modalities. However, this direct fusion is ineffective, since there exists a large gap due to the heterogeneous data representation between frame and event modalities. To address this issue, we explore a common-latent space as an intermediate bridge to mitigate the modality gap. In this work, we propose a novel common spatiotemporal fusion between frame and event modalities for high-dynamic scene optical flow, including visual boundary localization and motion correlation fusion. Specifically, in visual boundary localization, we figure out that frame and event share the similar spatiotemporal gradients, whose similarity distribution is consistent with the extracted boundary distribution. This motivates us to design the common spatiotemporal gradient to constrain the reference boundary localization. In motion correlation fusion, we discover that the frame-based motion possesses spatially dense but temporally discontinuous correlation, while the event-based motion has spatially sparse but temporally continuous correlation. This inspires us to use the reference boundary to guide the complementary motion knowledge fusion between the two modalities. Moreover, common spatiotemporal fusion can not only relieve the cross-modal feature discrepancy, but also make the fusion process interpretable for dense and continuous optical flow. Extensive experiments have been performed to verify the superiority of the proposed method.

光流估计事件相机多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。