arXiv:2507.08494cs.CV2025-07被引 2

用动态图网络统一建模多视角人体跟踪,无需预设轨迹片段。

One Graph to Track Them All: Dynamic GNNs for Single- and Multi-View Tracking

  • 构建时空动态图,融合空间、上下文与时间信息
  • 在新数据集上达到顶尖性能,尤其在遮挡场景表现优异
  • 适合研究多视角人体跟踪与复杂场景建模的学者

本文提出一种统一的可微模型,用于多人群体跟踪,无需依赖预计算的轨迹片段。模型构建动态时空图,聚合空间、上下文与时间信息,实现序列内信息的无缝传播。为提升遮挡处理能力,图结构还能编码场景特异性信息。研究引入一个包含25个部分重叠视角的大规模新数据集,具备精细场景重建和大量遮挡。实验表明,该模型在公开基准和新数据集上均达到当前最优性能,且对多种条件具有强适应性。数据集与方法将开源,以推动多人群体跟踪研究。

原文摘要 · Abstract (English)

This work presents a unified, fully differentiable model for multi-people tracking that learns to associate detections into trajectories without relying on pre-computed tracklets. The model builds a dynamic spatiotemporal graph that aggregates spatial, contextual, and temporal information, enabling seamless information propagation across entire sequences. To improve occlusion handling, the graph can also encode scene-specific information. We also introduce a new large-scale dataset with 25 partially overlapping views, detailed scene reconstructions, and extensive occlusions. Experiments show the model achieves state-of-the-art performance on public benchmarks and the new dataset, with flexibility across diverse conditions. Both the dataset and approach will be publicly released to advance research in multi-people tracking.

多视角跟踪动态图网络人体追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。