arXiv:2508.03749cs.CVeess.IV2025-08被引 1

用地铁站摄像头视频精准估算客流,提升安全与运营效率。

Closed-Circuit Television Data as an Emergent Data Source for Urban Rail Platform Crowding Estimation

  • 通过三种计算机视觉方法从监控画面提取人群特征。
  • 在600小时隐私保护数据上实现高精度实时客流估计。
  • 适合城市交通管理、智能调度与应急响应团队参考。

准确估算城市轨道交通站台客流可提升交通机构的运营决策能力,改善安全性、效率与乘客体验,尤其在应对拥挤问题时。然而,实时感知客流仍具挑战,常依赖闸机数据或人工观察等间接指标。近年来,闭路电视(CCTV)视频逐渐成为有潜力的数据源,可提供高精度、实时的客流估计。本研究对比了三种前沿计算机视觉方法:(a) 使用YOLOv11、RT-DETRv2和APGCC进行目标检测与计数;(b) 基于自训练视觉变压器Crowd-ViT的群体分类;(c) 采用DeepLabV3进行语义分割。此外,提出一种新型高效线性优化方法,从分割图中提取人数,并考虑图像中物体深度信息以反映乘客沿站台的分布情况。实验基于与华盛顿都会区交通局(WMATA)合作构建的隐私保护数据集,涵盖超过600小时视频。结果表明,计算机视觉方法能为客流估计提供实质性价值。该研究证明,仅依靠站台监控视频,即可实现更精确的实时拥挤度估计,进而支持及时的运营干预措施。

原文摘要 · Abstract (English)

Accurately estimating urban rail platform occupancy can enhance transit agencies' ability to make informed operational decisions, thereby improving safety, operational efficiency, and customer experience, particularly in the context of crowding. However, sensing real-time crowding remains challenging and often depends on indirect proxies such as automatic fare collection data or staff observations. Recently, Closed-Circuit Television (CCTV) footage has emerged as a promising data source with the potential to yield accurate, real-time occupancy estimates. The presented study investigates this potential by comparing three state-of-the-art computer vision approaches for extracting crowd-related features from platform CCTV imagery: (a) object detection and counting using YOLOv11, RT-DETRv2, and APGCC; (b) crowd-level classification via a custom-trained Vision Transformer, Crowd-ViT; and (c) semantic segmentation using DeepLabV3. Additionally, we present a novel, highly efficient linear-optimization-based approach to extract counts from the generated segmentation maps while accounting for image object depth and, thus, for passenger dispersion along a platform. Tested on a privacy-preserving dataset created in collaboration with the Washington Metropolitan Area Transit Authority (WMATA) that encompasses more than 600 hours of video material, our results demonstrate that computer vision approaches can provide substantive value for crowd estimation. This work demonstrates that CCTV image data, independent of other data sources available to a transit agency, can enable more precise real-time crowding estimation and, eventually, timely operational responses for platform crowding mitigation.

客流估计计算机视觉智慧交通视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。