arXiv:2508.17205cs.CVcs.AI2025-08被引 4

用多智能体框架提升高速公路场景理解,兼顾精度与效率。

Multi-Agent Visual-Language Reasoning for Comprehensive Highway Scene Understanding

  • 大模型生成任务提示,小模型高效推理多模态视频
  • 在多种天气和路况下同时完成天气、路面湿滑、拥堵检测
  • 可部署于老旧摄像头系统,适合高风险路段实时预警

本文提出一种多智能体框架,用于全面理解高速公路场景。该框架基于专家混合策略,将通用视觉语言模型(如GPT-4o)结合领域知识,生成任务特定的思维链(CoT)提示,进而指导小型高效模型(如Qwen2.5-VL-7B)对短时视频及互补模态进行推理。该框架同步解决天气分类、路面湿滑评估与交通拥堵检测等关键感知任务,在保证准确率的同时实现计算效率平衡。为验证效果,我们构建了三个专用数据集,其中路面湿滑数据集融合视频流与道路气象传感器数据,体现多模态推理优势。实验表明,系统在多样交通与环境条件下表现稳定。从部署角度看,框架可无缝集成现有交通监控系统,并适用于急弯、易涝低地、结冰桥梁等高风险农村路段,实现持续监测与及时预警,即使在资源受限环境中亦可运行。

原文摘要 · Abstract (English)

This paper introduces a multi-agent framework for comprehensive highway scene understanding, designed around a mixture-of-experts strategy. In this framework, a large generic vision-language model (VLM), such as GPT-4o, is contextualized with domain knowledge to generates task-specific chain-of-thought (CoT) prompts. These fine-grained prompts are then used to guide a smaller, efficient VLM (e.g., Qwen2.5-VL-7B) in reasoning over short videos, along with complementary modalities as applicable. The framework simultaneously addresses multiple critical perception tasks, including weather classification, pavement wetness assessment, and traffic congestion detection, achieving robust multi-task reasoning while balancing accuracy and computational efficiency. To support empirical validation, we curated three specialized datasets aligned with these tasks. Notably, the pavement wetness dataset is multimodal, combining video streams with road weather sensor data, highlighting the benefits of multimodal reasoning. Experimental results demonstrate consistently strong performance across diverse traffic and environmental conditions. From a deployment perspective, the framework can be readily integrated with existing traffic camera systems and strategically applied to high-risk rural locations, such as sharp curves, flood-prone lowlands, or icy bridges. By continuously monitoring the targeted sites, the system enhances situational awareness and delivers timely alerts, even in resource-constrained environments.

多智能体视觉语言交通监控多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。