arXiv:2604.15312cs.CV2026-04

双向跨模态提示提升动态场景立体感知精度

Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo

论文配图:Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo
图 1 · 摘自论文原文
  • 双向融合事件与帧图像特征,增强跨模态匹配
  • 在快速运动和复杂光照下实现更优3D感知性能
  • 适合高动态场景的视觉系统研究者使用

传统帧相机在动态场景中受限于有限的时间分辨率和运动模糊。事件相机提供更高动态范围的视觉表征,避免此类问题。两者互补特性使事件-帧异构立体视觉在快速运动和挑战性光照条件下具有可靠的3D感知潜力。然而,模态差异常导致领域特有线索被弱化,影响跨模态匹配。本文提出Bi-CMPStereo,一种新型双向跨模态提示框架,充分挖掘双域语义与结构特征以实现鲁棒匹配。方法在目标规范空间内学习精细对齐的立体表示,并通过将每种模态投影至事件与帧域来融合互补信息。大量实验表明,该方法在准确率与泛化能力上显著优于现有最先进方法。

原文摘要 · Abstract (English)

Conventional frame-based cameras capture rich contextual information but suffer from limited temporal resolution and motion blur in dynamic scenes. Event cameras offer an alternative visual representation with higher dynamic range free from such limitations. The complementary characteristics of the two modalities make event-frame asymmetric stereo promising for reliable 3D perception under fast motion and challenging illumination. However, the modality gap often leads to marginalization of domain-specific cues essential for cross-modal stereo matching. In this paper, we introduce Bi-CMPStereo, a novel bidirectional cross-modal prompting framework that fully exploits semantic and structural features from both domains for robust matching. Our approach learns finely aligned stereo representations within a target canonical space and integrates complementary representations by projecting each modality into both event and frame domains. Extensive experiments demonstrate that our approach significantly outperforms state-of-the-art methods in accuracy and generalization.

立体视觉事件相机跨模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。