arXiv:2410.21743cs.CVcs.RO2024-10中稿 · WACV 2025被引 1

提出无中介跨模态特征提取框架,提升事件相机与图像数据匹配精度

EI-Nexus: Towards Unmediated and Flexible Inter-Modality Local Feature Extraction and Matching for Event-Image Data

  • 通过局部特征蒸馏实现视角不变的跨模态关键点提取
  • 引入上下文聚合使特征匹配性能显著提升,达到新基准最优
  • 构建首个跨模态位姿估计基准,适合多模态感知研究者使用

事件相机具有高时间分辨率和高动态范围,但其在事件-图像数据跨模态局部特征提取与匹配方面研究有限。本文提出EI-Nexus框架,无需中间转换,灵活集成两种模态专用的关键点提取器与特征匹配器。为应对视点和模态变化,提出局部特征蒸馏(LFD),将图像提取器中已学习的视点一致性迁移至事件提取器,确保特征对应鲁棒性。同时,借助上下文聚合(CA)机制,显著提升特征匹配效果。进一步建立首个跨模态特征匹配基准MVSEC-RPE和EC-RPE,用于评估事件-图像数据上的相对位姿估计。所提方法优于依赖显式模态转换的传统方法,实现更直接、灵活的特征提取与匹配,在MVSEC-RPE和EC-RPE基准上取得当前最佳性能。代码与基准数据集将在https://github.com/ZhonghuaYi/EI-Nexus_official公开。

原文摘要 · Abstract (English)

Event cameras, with high temporal resolution and high dynamic range, have limited research on the inter-modality local feature extraction and matching of event-image data. We propose EI-Nexus, an unmediated and flexible framework that integrates two modality-specific keypoint extractors and a feature matcher. To achieve keypoint extraction across viewpoint and modality changes, we bring Local Feature Distillation (LFD), which transfers the viewpoint consistency from a well-learned image extractor to the event extractor, ensuring robust feature correspondence. Furthermore, with the help of Context Aggregation (CA), a remarkable enhancement is observed in feature matching. We further establish the first two inter-modality feature matching benchmarks, MVSEC-RPE and EC-RPE, to assess relative pose estimation on event-image data. Our approach outperforms traditional methods that rely on explicit modal transformation, offering more unmediated and adaptable feature extraction and matching, achieving better keypoint similarity and state-of-the-art results on the MVSEC-RPE and EC-RPE benchmarks. The source code and benchmarks will be made publicly available at https://github.com/ZhonghuaYi/EI-Nexus_official.

跨模态事件相机特征匹配位姿估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。