用运动引导注意力与经验分解结合,精准识别道路拥堵状态。
Hybrid Congestion Classification Framework Using Flow-Guided Attention and Empirical Mode Decomposition

- 用光流引导空间和通道注意力,聚焦动态区域特征。
- 通过EMD分解流量统计,提取内在时序成分,准确率达97.5%。
- 适合交通监控、智能驾驶等需实时拥堵判断的场景。
精准的交通拥堵分类需同时捕捉道路场景上下文与非平稳的交通运动特性,但现有方法多孤立处理二者。视觉方法依赖外观线索并使用标准时序池化,易受静态基础设施干扰;信号方法虽能刻画时序动态,却缺乏场景级空间定位能力。为此,本文提出FLO-EMD,融合运动引导注意力与经验性数据驱动时序分解。密集光流用于指导通道与空间注意力,使RGB特征聚焦于运动相关区域;同时,聚合的光流统计形成紧凑的运动轨迹,并通过经验模态分解(EMD)提取内在时序分量。最终将EMD嵌入与学习到的时空表示融合,实现轻度、中度、重度拥堵的分类。在四个监控网络的1,050段五秒视频上测试,整体测试准确率达97.5%(加权F1=0.9742),优于主流基线,且在多种环境条件下保持鲁棒;消融与敏感性分析进一步量化了EMD、固有模态函数数量及运动描述子的贡献。
原文摘要 · Abstract (English)
Accurate traffic congestion classification requires models that jointly capture roadway scene context and non-stationary traffic motion, yet most prior work treats these requirements in isolation. Vision-based methods often depend on appearance cues with standard temporal pooling, which can bias predictions toward static infrastructure, whereas signal-based approaches characterize temporal dynamics but lack the spatial context needed for scene-level localization. These complementary limitations motivate a unified framework that links motion evidence to spatial feature selection while preserving data-adaptive temporal characterization. This study therefore proposes FLO-EMD, a hybrid approach that couples motion-guided attention with empirical, data-driven temporal decomposition. Dense optical flow guides channel and spatial attention so that RGB features are refined toward motion-relevant regions. In parallel, aggregated flow statistics form compact motion traces that are decomposed using Empirical Mode Decomposition (EMD) to extract intrinsic temporal components. The resulting EMD embedding is fused with learned spatiotemporal representations to classify light, medium, and heavy congestion. Experiments on 1,050 five-second clips from four surveillance networks show that FLO-EMD achieves 97.5% overall test accuracy (weighted F1 = 0.9742), outperforming established baselines and remaining robust across diverse environmental conditions; ablation and sensitivity analyses further quantify the contributions of EMD, the number of intrinsic mode functions, and the selected motion descriptors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。