arXiv:2602.12919cs.CVcs.AI2026-02

构建首个事件流视觉定位高质量基准数据集,支持高动态、低光照场景下的精准定位。

EPRBench: A High-Quality Benchmark Dataset for Event Stream Based Visual Place Recognition

  • 采集10K事件序列与6.5万帧事件图像,覆盖手持与车载多种真实场景。
  • 15种先进算法在该数据集上基准测试,新融合框架实现更高精度与可解释性。
  • 首次引入大模型生成语义描述,助力事件流与文本跨模态融合研究。

基于事件流的视觉定位(VPR)是新兴研究方向,能有效应对传统可见光相机在低光照、过曝及高速运动等挑战性条件下的不稳定性。针对该领域专用数据集匮乏的问题,本文提出EPRBench,一个专为事件流视觉定位设计的高质量基准数据集。EPRBench包含10,000个事件序列和65,000帧事件图像,通过手持与车载两种采集方式,全面覆盖多视角、多天气和多光照条件下的真实挑战。为支持语义感知与语言融合的VPR研究,我们利用大语言模型(LLM)生成场景描述,并经人工标注优化,为事件流感知流程中集成大模型奠定基础。为促进系统化评估,我们在EPRBench上实现并基准测试了15种先进VPR算法,提供可靠对比基线。此外,我们提出一种新型多模态融合范式:利用LLM从原始事件流生成文本场景描述,进而指导空间注意力的标记选择、跨模态特征融合与多尺度表征学习。该框架不仅实现高精度定位,还生成可解释的推理过程,显著提升模型透明度与可解释性。数据集与源代码将公开于https://github.com/Event-AHU/Neuromorphic_ReID。

原文摘要 · Abstract (English)

Event stream-based Visual Place Recognition (VPR) is an emerging research direction that offers a compelling solution to the instability of conventional visible-light cameras under challenging conditions such as low illumination, overexposure, and high-speed motion. Recognizing the current scarcity of dedicated datasets in this domain, we introduce EPRBench, a high-quality benchmark specifically designed for event stream-based VPR. EPRBench comprises 10K event sequences and 65K event frames, collected using both handheld and vehicle-mounted setups to comprehensively capture real-world challenges across diverse viewpoints, weather conditions, and lighting scenarios. To support semantic-aware and language-integrated VPR research, we provide LLM-generated scene descriptions, subsequently refined through human annotation, establishing a solid foundation for integrating LLMs into event-based perception pipelines. To facilitate systematic evaluation, we implement and benchmark 15 state-of-the-art VPR algorithms on EPRBench, offering a strong baseline for future algorithmic comparisons. Furthermore, we propose a novel multi-modal fusion paradigm for VPR: leveraging LLMs to generate textual scene descriptions from raw event streams, which then guide spatially attentive token selection, cross-modal feature fusion, and multi-scale representation learning. This framework not only achieves highly accurate place recognition but also produces interpretable reasoning processes alongside its predictions, significantly enhancing model transparency and explainability. The dataset and source code will be released on https://github.com/Event-AHU/Neuromorphic_ReID

事件流视觉定位多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。