arXiv:2510.10342cs.CV2025-10

用视觉语言和运动分析,给交通拥堵分级,更准更可解释。

Ordinal Scale Traffic Congestion Classification with Multi-Modal Vision-Language and Motion Analysis

  • 融合视觉语言模型与运动分析,识别拥堵等级
  • 在1到5级的有序分类中准确率达76.7%
  • 适合智能交通系统开发与城市管理者使用

精准的交通拥堵分类对智能交通系统和实时城市交通管理至关重要。本文提出一种多模态框架,结合开放词汇视觉语言推理(CLIP)、目标检测(YOLO-World)以及基于MOG2背景分割的运动分析。该系统将拥堵水平预测为从1(自由流)到5(严重拥堵)的有序等级,实现语义对齐和时间一致性。为提升可解释性,引入基于运动的置信度加权并生成带注释的可视化输出。实验结果表明,该模型准确率达76.7%,F1分数为0.752,加权卡帕系数(QWK)达0.684,显著优于单模态基线。结果验证了该框架在保持有序结构及融合视觉语言与运动模态方面的有效性。未来工作包括引入车辆尺寸与精细化密度指标。

原文摘要 · Abstract (English)

Accurate traffic congestion classification is essential for intelligent transportation systems and real-time urban traffic management. This paper presents a multimodal framework combining open-vocabulary visual-language reasoning (CLIP), object detection (YOLO-World), and motion analysis via MOG2-based background subtraction. The system predicts congestion levels on an ordinal scale from 1 (free flow) to 5 (severe congestion), enabling semantically aligned and temporally consistent classification. To enhance interpretability, we incorporate motion-based confidence weighting and generate annotated visual outputs. Experimental results show the model achieves 76.7 percent accuracy, an F1 score of 0.752, and a Quadratic Weighted Kappa (QWK) of 0.684, significantly outperforming unimodal baselines. These results demonstrate the framework's effectiveness in preserving ordinal structure and leveraging visual-language and motion modalities. Future enhancements include incorporating vehicle sizing and refined density metrics.

交通拥堵多模态视觉语言运动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。