arXiv:2604.19318cs.CV2026-04

用视觉与地面交互的Transformer提升大场景人群追踪效果

Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes

论文配图:Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
图 1 · 摘自论文原文
  • 引入视图与地面平面的交互机制,增强多视角追踪
  • 在两个新构建的大场景数据集上显著超越现有方法
  • 适合关注真实复杂场景下人群追踪的研究者

多视角人群追踪旨在估计场景地面上每个人的轨迹。当前主流方法依赖基于CNN的多视角追踪架构,且多数在较小数据集(如Wildtrack和MultiviewX)上评估,这些数据集场景小、帧数少(仅数十帧),难以反映真实世界中复杂的场景规模与遮挡情况。本文提出一种基于Transformer的多视角人群追踪模型MVTrackTrans,通过引入相机视角与地面平面间的交互来提升追踪性能。同时,为更好评估,我们收集并标注了两个大规模真实场景多视角追踪数据集MVCrowdTrack和CityTrack,覆盖更大场景范围和更长时间跨度。在两个新数据集上,所提模型相比现有方法表现更优,验证了模型设计在处理大场景时的优势。代码与数据集已公开于https://github.com/zqyq/MVTrackTrans。

原文摘要 · Abstract (English)

Multi-view crowd tracking estimates each person's tracking trajectories on the ground of the scene. Recent research works mainly rely on CNNs-based multi-view crowd tracking architectures, and most of them are evaluated and compared on relatively small datasets, such as Wildtrack and MultiviewX. Since these two datasets are collected in small scenes and only contain tens of frames in the evaluation stage, it is difficult for the current methods to be applied to real-world applications where scene size and occlusion are more complicated. In this paper, we propose a Transformer-based multi-view crowd tracking model, \textit{MVTrackTrans}, which adopts interactions between camera views and the ground plane for enhanced multi-view tracking performance. Besides, for better evaluation, we collect and label two large real-world multi-view tracking datasets, MVCrowdTrack and CityTrack, which contain a much larger scene size over a longer time period. Compared with existing methods on the two large and new datasets, the proposed MVTrackTrans model achieves better performance, demonstrating the advantages of the model design in dealing with large scenes. We believe the proposed datasets and model will push the frontiers of the task to more practical scenarios, and the datasets and code are available at: https://github.com/zqyq/MVTrackTrans.

人群追踪Transformer多视角大场景

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。