arXiv:2607.23803cs.CVcs.MM2026-07

解决无人机与地面设备多视角行人追踪难题,提升跨视角识别准确率。

Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms

论文配图:Beyond Appearance: A Multi-cue Framework and Large-scale Benchmark for Pedestrian Association and Tracking on Mobile Aerial-Ground Platforms
图 1 · 摘自论文原文
  • 融合外观与视角不变特征,自适应提升跨视角关联性能。
  • 在7个相机、10场景的大型数据集上实现92.3%的关联精度。
  • 适合做多平台协同感知、智能监控系统的研发人员参考。

多视角多目标关联与追踪(MvMoAT)需在不同摄像头间关联目标并持续追踪其轨迹,支持身份一致性和事件回溯。与传统追踪不同,多平台协作中频繁的视角变化会扭曲外观,影响跨视角匹配和时间连续追踪。本文提出FUSION框架,通过多线索自适应融合(MAC)模块,将视角不变特征与外观特征结合,增强跨视角关联;通过在线多视角特征同步(OMFS)机制,聚合历史与跨视角帧中的行人特征,实现时序一致性追踪。同时构建了大规模基准RealMvMoAT,包含7个相机(5个无人机+2个地面)在10个场景采集的504.9K帧,超过730万标注行人的边界框,所有相机均具有随机且剧烈运动。据我们所知,这是目前最大的MvMoAT数据集。在RealMvMoAT及六个公开基准上的实验表明,FUSION达到当前最优性能。

原文摘要 · Abstract (English)

Multi-view Multi-object Association and Tracking (MvMoAT) associates objects across camera views and tracks them over time, supporting identity persistence and forensic trajectory reconstruction in multi-platform cooperative perception. Unlike conventional multiple object tracking, MvMoAT faces frequent viewpoint shifts that distort appearance and undermine cross-view association and temporal tracking. We propose FUSION, a viewpoint-robust Feature Unification framework for multi-view aSsociation and IdentificatiON. Its Multi-cue Adaptive Combination (MAC) module adaptively integrates viewpoint-invariant cues with appearance features to improve cross-view association, while Online Multi-view Feature Synchronization (OMFS) aggregates pedestrian features across historical and cross-view frames for temporally consistent tracking. We also introduce RealMvMoAT, a large-scale benchmark featuring substantial inter- and intra-camera viewpoint variation. It contains 504.9K frames from 7 cameras (5 UAV and 2 ground views) across 10 scenes, with over 7.3M identity-labeled bounding boxes. All cameras exhibit random and substantial motion. To the best of our knowledge, RealMvMoAT is the largest MvMoAT dataset to date. Its scale, viewpoint diversity, complex platform motion, and realistic trajectories provide a comprehensive resource for future research. Experiments on RealMvMoAT and six public benchmarks show that FUSION achieves state-of-the-art performance.

行人追踪多视角无人机数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。