系统梳理视觉技术在交通监控中的应用与挑战,提出融合大模型的升级路径。
Vision Technologies with Applications in Traffic Surveillance Systems: A Holistic Survey
- 按感知层级分类解析目标检测、行为理解等任务方法
- 发现五类核心局限:数据退化、学习受限、语义不足等
- 推荐大模型、协同感知等前沿方案,适合智能交通研究者
交通监控系统(TSS)在现代智能交通中日益重要,视觉技术在场景感知与理解中起核心作用。现有综述多聚焦孤立环节,缺乏连接低层与高层感知任务的综合分析框架,尤其对新兴技术关注不足。本文系统回顾了TSS中的视觉技术,涵盖低层任务(目标检测、分类、跟踪)与高层任务(参数估计、异常检测、行为理解)。我们对每项任务进行方法分类与性能评估,揭示当前TSS存在五大根本局限:复杂场景下感知数据退化、数据驱动学习约束、语义理解缺口、感知覆盖不足及计算资源需求高。为应对挑战,系统分析五类现有与潜在趋势:先进感知增强、高效学习范式、知识增强理解、协同感知架构与高效计算框架,并评估其实际应用潜力。此外,评估基础模型在TSS中的变革性潜力,其展现出卓越的零样本学习能力、强泛化性与复杂任务推理能力。本综述构建统一分析框架,系统剖析现有瓶颈与解决方案,提出整合新兴技术(尤其是基础模型)的结构化路线图,以提升TSS能力。
原文摘要 · Abstract (English)
Traffic Surveillance Systems (TSS) have become increasingly crucial in modern intelligent transportation systems, with vision technologies playing a central role for scene perception and understanding. While existing surveys typically focus on isolated aspects of TSS, a comprehensive analytical framework bridging low-level and high-level perception tasks, particularly considering emerging technologies, remains lacking. This paper presents a systematic review of vision technologies in TSS, examining both low-level perception tasks (object detection, classification, and tracking) and high-level perception tasks (parameter estimation, anomaly detection, and behavior understanding). Specifically, we first provide a detailed methodological categorization and comprehensive performance evaluation for each task. Our investigation reveals five fundamental limitations in current TSS: perceptual data degradation in complex scenarios, data-driven learning constraints, semantic understanding gaps, sensing coverage limitations and computational resource demands. To address these challenges, we systematically analyze five categories of current approaches and potential trends: advanced perception enhancement, efficient learning paradigms, knowledge-enhanced understanding, cooperative sensing frameworks and efficient computing frameworks, critically assessing their real-world applicability. Furthermore, we evaluate the transformative potential of foundation models in TSS, which exhibit remarkable zero-shot learning abilities, strong generalization, and sophisticated reasoning capabilities across diverse tasks. This review provides a unified analytical framework bridging low-level and high-level perception tasks, systematically analyzes current limitations and solutions, and presents a structured roadmap for integrating emerging technologies, particularly foundation models, to enhance TSS capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。