arXiv:2505.12254cs.CVcs.AI2025-05被引 4

首个面向行人视角的多模态街景识别数据集与评测平台

MMS-VPR: Multimodal Street-Level Visual Place Recognition Dataset and Benchmark

  • 构建行人视角街景数据集,融合图像、视频与文本多模态信息
  • 覆盖成都7万平米商业区208个地点,含昼夜、多角度与七年社交媒体数据
  • 提供标准化评测框架,支持视觉、视频、文本模态协同分析

现有视觉场景识别(VPR)数据集多基于车载影像,缺乏多模态多样性,且对非西方城市密集人行街景代表性不足。本文提出MMS-VPR,一个大规模行人专用街景识别多模态数据集。该数据集包含110,529张图像和2,527段视频剪辑,覆盖成都约70,800平方米开放商业区内的208个位置,实地采集于2024年,社交媒体数据跨度为2019–2025年,兼具细粒度时间分辨率与长期时间覆盖。每个地点均包含昼夜交替、多视角及多模态标注(含GPS坐标、时间戳与语义文本元数据)。同时发布MMS-VPRlib,一个统一基准平台,整合常用VPR数据集与前沿方法,采用标准化可复现流程。平台模块化支持数据预处理、多模态建模(CNN/RNN/Transformer)、信号增强、特征对齐、融合与评估。该平台突破传统仅图像范式,系统性挖掘视觉、视频与文本的互补性。数据集链接:https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR;基准代码:https://github.com/yiasun/MMS-VPRlib。

原文摘要 · Abstract (English)

Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and underrepresent dense pedestrian street scenes, particularly in non-Western urban contexts. We introduce MMS-VPR, a large-scale multimodal dataset for street-level place recognition in pedestrian-only environments. MMS-VPR comprises 110,529 images and 2,527 video clips across 208 locations in a ~70,800 $m^2$ open-air commercial district in Chengdu, China. Field data were collected in 2024, while social media data span seven years (2019-2025), providing both fine-grained temporal granularity and long-term temporal coverage. Each location features comprehensive day-night coverage, multiple viewing angles, and multimodal annotations including GPS coordinates, timestamps, and semantic textual metadata. We further release MMS-VPRlib, a unified benchmarking platform that consolidates commonly used VPR datasets and state-of-the-art methods under a standardized, reproducible pipeline. MMS-VPRlib provides modular components for data pre-processing, multimodal modeling (CNN/RNN/Transformer), signal enhancement, alignment, fusion, and performance evaluation. This platform moves beyond traditional image-only paradigms, enabling systematic exploitation of complementary visual, video, and textual modalities. The dataset is available at https://huggingface.co/datasets/Yiwei-Ou/MMS-VPR and the benchmark at https://github.com/yiasun/MMS-VPRlib.

多模态街景识别数据集行人视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。