arXiv:2609.07685cs.CV2026-09

构建真实密集人群轨迹预测基准,评估从检测到预测的全流程性能。

CrowdTraj: A Benchmark for Dense Crowd Trajectory Prediction in Realistic Crowded Environments

论文配图:CrowdTraj: A Benchmark for Dense Crowd Trajectory Prediction in Realistic Crowded Environments
图 1 · 摘自论文原文
  • 基于监控视角与严重遮挡,支持端到端评估跟踪与预测。
  • 单场景平均1146人,帧密度达372人,标注超320万头框。
  • 揭示现有模型在跟踪噪声和计算效率上的实际短板。

在真实应用中,行人轨迹预测依赖于检测与跟踪系统的输入。以往的轨迹预测基准要么行人交互稀疏,要么假设追踪输入完美,或依赖俯视图以减少遮挡和透视失真,限制了对真实密集场景的评估。本文提出CrowdTraj,一个面向自然密集人群场景的行人轨迹预测基准。不同于以往数据集,CrowdTraj支持在闭路电视视角下,于严重遮挡条件下,从检测经由跟踪到轨迹预测的端到端评估。它还捕捉多样且自然的行人行为,包括现有基准罕见的突发方向变化。CrowdTraj包含五个多样化场景,每场景平均有1,146名唯一行人,帧级最大密度范围为114至372人,共标注超过320万个人头边界框。通过各场景的单应性矩阵提供像素坐标与真实世界坐标,支持物理意义下的分析。实验结果显示,在最密集场景中,追踪精度(IDF1)下降至0.68至0.70,相较较稀疏场景的约0.90显著降低。轨迹预测训练在密集场景中也大幅增加计算成本,训练时间最高提升8倍。这些发现表明,CrowdTraj暴露了当前轨迹预测流程在鲁棒性与计算可扩展性方面的局限,而这些局限在现有稀疏人群基准中难以察觉。

原文摘要 · Abstract (English)

In real-world applications, pedestrian trajectory prediction models rely on inputs from detection and tracking systems. Prior trajectory prediction benchmarks either contain relatively sparse pedestrian interactions, assume perfect tracking inputs, or rely on overhead viewpoints that minimize occlusion and perspective distortion, limiting evaluation in realistic dense-crowd scenarios. We present CrowdTraj, a benchmark for pedestrian trajectory prediction in natural dense crowd scenes. Unlike previous datasets, CrowdTraj supports end-to-end evaluation from detection through tracking to trajectory prediction under severe occlusion in CCTV views. It also captures diverse, natural pedestrian behaviours, including abrupt directional changes rarely observed in existing benchmarks. CrowdTraj includes five diverse scenes, with an average of 1,146 unique pedestrians per scene, maximum frame-level densities ranging from 114 to 372 pedestrians, and over 3.2 million annotated head bounding boxes. CrowdTraj provides pixel and real-world coordinates via per-scene homography matrices for physically meaningful analysis. Our experimental results show that tracking accuracy (IDF1) drops to 0.68 to 0.70 in the densest scenes, compared with approximately 0.90 in less crowded scenes. Trajectory prediction training also becomes substantially more computationally expensive in dense scenes, with training times increasing by up to 8 times. These findings show that CrowdTraj exposes limitations in current trajectory prediction pipelines that remain hidden on existing sparse-crowd benchmarks, particularly in robustness to tracking noise and computational scalability.

轨迹预测密集人群多模态基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。