arXiv:2510.09981cs.CVeess.IV2025-10

用AI和语言模型分析交通摄像头,实现大规模实时交通监测。

Scaling Traffic Insights with AI and Language Model-Powered Camera Systems for Data-Driven Transportation Decision Making

  • 用微调的YOLOv11实时提取交通密度与分类数据
  • 非固定摄像头视角下仍保持9%的车辆减少检测精度
  • 适合城市交通政策评估与应急响应决策者

准确、可扩展的交通监控对实时与长期交通管理至关重要,尤其在自然灾害、大型施工或政策变化(如纽约市首例拥堵收费)期间。然而,由于安装、维护和数据管理成本高,传感器部署受限。交通摄像头虽具成本优势,但现有视频分析难以应对动态视角和大规模数据。本研究提出一个基于AI的端到端框架,利用现有摄像头基础设施实现高分辨率、长期性的大规模分析。采用在本地城市场景训练的微调YOLOv11模型,实时提取多模态交通密度与分类指标。针对非固定云台镜头导致的视角不一致,提出一种图结构视角归一化方法。同时引入领域特定大语言模型,处理24/7视频流数据,自动生成频繁的交通模式摘要,远超人工处理能力。系统在2025年纽约拥堵收费早期推行阶段验证,使用约1000个摄像头采集超过900万张图像。结果显示,拥堵缓解区工作日私家车密度下降9%,货车流量提前减少并出现反弹迹象,人行道与自行车道活动量持续上升。实验表明,示例提示显著提升LLM数值准确性,减少幻觉。该框架展现出无需人工干预即可支持大规模、政策相关交通监测的实用潜力。

原文摘要 · Abstract (English)

Accurate, scalable traffic monitoring is critical for real-time and long-term transportation management, particularly during disruptions such as natural disasters, large construction projects, or major policy changes like New York City's first-in-the-nation congestion pricing program. However, widespread sensor deployment remains limited due to high installation, maintenance, and data management costs. While traffic cameras offer a cost-effective alternative, existing video analytics struggle with dynamic camera viewpoints and massive data volumes from large camera networks. This study presents an end-to-end AI-based framework leveraging existing traffic camera infrastructure for high-resolution, longitudinal analysis at scale. A fine-tuned YOLOv11 model, trained on localized urban scenes, extracts multimodal traffic density and classification metrics in real time. To address inconsistencies from non-stationary pan-tilt-zoom cameras, we introduce a novel graph-based viewpoint normalization method. A domain-specific large language model was also integrated to process massive data from a 24/7 video stream to generate frequent, automated summaries of evolving traffic patterns, a task far exceeding manual capabilities. We validated the system using over 9 million images from roughly 1,000 traffic cameras during the early rollout of NYC congestion pricing in 2025. Results show a 9% decline in weekday passenger vehicle density within the Congestion Relief Zone, early truck volume reductions with signs of rebound, and consistent increases in pedestrian and cyclist activity at corridor and zonal scales. Experiments showed that example-based prompts improved LLM's numerical accuracy and reduced hallucinations. These findings demonstrate the framework's potential as a practical, infrastructure-ready solution for large-scale, policy-relevant traffic monitoring with minimal human intervention.

交通监控AI分析大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。