arXiv:2503.08121cs.CV2025-03CVPR被引 22

构建首个空地跨平台视频行人重识别大基准,挑战视角、尺度和分辨率差异。

AG-VPReID: A Challenging Large-Scale Benchmark for Aerial-Ground Video-based Person Re-Identification

  • 设计三流融合框架,分别应对运动不一致、外观变化和尺度差异。
  • 涵盖6632人、960万帧,无人机高度15-120米,真实复杂场景。
  • 适合关注空地协同监控、多视角行人识别的研究者使用。

我们提出AG-VPReID,一个面向空地跨平台视频行人重识别的大规模数据集,包含6632名个体、32321个轨迹和超过960万帧图像,由无人机(15-120米高度)、CCTV及可穿戴相机采集。该数据集为评估在显著视角变化、尺度波动与分辨率差异下的鲁棒性提供了真实世界基准。为此,我们设计了端到端的AG-VPReID-Net框架,包含三个互补分支:(1) 自适应时空流处理运动模式不一致,促进时序特征学习;(2) 归一化外观流利用物理启发技术缓解分辨率与外观变化;(3) 多尺度注意力流应对无人机不同高度带来的尺度差异。融合各流的视觉-语义线索,生成鲁棒且视角不变的全身表征。大量实验表明,AG-VPReID-Net在新数据集及现有视频ReID基准上均优于当前最优方法,展现其有效性与泛化能力。然而,所有方法在AG-VPReID上的性能差距仍显著,凸显数据集的挑战性。数据集、代码与训练模型已开源于https://github.com/agvpreid25/AG-VPReID-Net。

原文摘要 · Abstract (English)

We introduce AG-VPReID, a new large-scale dataset for aerial-ground video-based person re-identification (ReID) that comprises 6,632 subjects, 32,321 tracklets and over 9.6 million frames captured by drones (altitudes ranging from 15-120m), CCTV, and wearable cameras. This dataset offers a real-world benchmark for evaluating the robustness to significant viewpoint changes, scale variations, and resolution differences in cross-platform aerial-ground settings. In addition, to address these challenges, we propose AG-VPReID-Net, an end-to-end framework composed of three complementary streams: (1) an Adapted Temporal-Spatial Stream addressing motion pattern inconsistencies and facilitating temporal feature learning, (2) a Normalized Appearance Stream leveraging physics-informed techniques to tackle resolution and appearance changes, and (3) a Multi-Scale Attention Stream handling scale variations across drone altitudes. We integrate visual-semantic cues from all streams to form a robust, viewpoint-invariant whole-body representation. Extensive experiments demonstrate that AG-VPReID-Net outperforms state-of-the-art approaches on both our new dataset and existing video-based ReID benchmarks, showcasing its effectiveness and generalizability. Nevertheless, the performance gap observed on AG-VPReID across all methods underscores the dataset's challenging nature. The dataset, code and trained models are available at https://github.com/agvpreid25/AG-VPReID-Net.

行人重识别空地协同视频分析多视角

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。