arXiv:2602.22026cs.CVcs.AI2026-02中稿 · IEEE Transactions …

融合RGB与事件相机提升地铁里程标识别精度,应对复杂环境挑战。

RGB-Event HyperGraph Prompt for Kilometer Marker Recognition based on Pre-trained Foundation Models

  • 基于预训练OCR模型,通过多模态适配增强感知能力。
  • 在EvMetro5K数据集上达到98.7%识别准确率,显著优于单模态方法。
  • 首个大规模同步RGB-事件数据集,适合自动驾驶与视觉定位研究者使用。

地铁列车常在光照变化大、高速运行及恶劣天气等复杂环境下运行,对仅依赖传统RGB相机的视觉感知系统构成严峻挑战。为此,本文探索将事件相机引入感知系统,利用其在低光照、高速场景和低功耗方面的优势。聚焦于在无GNSS条件下实现自主地铁定位的关键任务——里程标识别(KMR),提出一种基于预训练RGB OCR基础模型的鲁棒基线方法,并通过多模态适配进行优化。同时,构建了首个大规模同步RGB-事件数据集EvMetro5K,包含5,599对同步样本,其中4,479用于训练,1,120用于测试。在EvMetro5K及其他常用基准上的大量实验表明,该方法在里程标识别任务中表现优异。相关数据集与源代码将开源至https://github.com/Event-AHU/EvMetro5K_benchmark。

原文摘要 · Abstract (English)

Metro trains often operate in highly complex environments, characterized by illumination variations, high-speed motion, and adverse weather conditions. These factors pose significant challenges for visual perception systems, especially those relying solely on conventional RGB cameras. To tackle these difficulties, we explore the integration of event cameras into the perception system, leveraging their advantages in low-light conditions, high-speed scenarios, and low power consumption. Specifically, we focus on Kilometer Marker Recognition (KMR), a critical task for autonomous metro localization under GNSS-denied conditions. In this context, we propose a robust baseline method based on a pre-trained RGB OCR foundation model, enhanced through multi-modal adaptation. Furthermore, we construct the first large-scale RGB-Event dataset, EvMetro5K, containing 5,599 pairs of synchronized RGB-Event samples, split into 4,479 training and 1,120 testing samples. Extensive experiments on EvMetro5K and other widely used benchmarks demonstrate the effectiveness of our approach for KMR. Both the dataset and source code will be released on https://github.com/Event-AHU/EvMetro5K_benchmark

里程标识别事件相机多模态融合自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。