arXiv:2502.11304cs.AIcs.CL2025-02被引 7

用多模态大模型+实例分割提升交通监控精度与效率

Leveraging Multimodal-LLMs Assisted by Instance Segmentation for Intelligent Traffic Monitoring

  • 结合多模态大模型与实例分割,实时分析城市交通场景
  • 车辆定位准确率达84.3%,方向判断准确率76.4%
  • 适用于智慧交通系统,适合城市管理者与算法研究者

智能交通系统(ITS)需要高效可靠的交通监控系统,通过传感器和摄像头追踪车辆动态、优化交通流、缓解拥堵、提升道路安全并实现实时自适应控制。本研究在真实Quanser交互式实验平台中,利用LLaVA视觉定位多模态大语言模型(LLM)完成交通监控任务,涵盖交叉口、拥堵与碰撞等场景。多个城市位置的摄像头实时采集仿真图像,输入至LLaVA模型进行查询分析。集成的实例分割模型突出显示车辆与行人等关键元素,提升训练效果与处理吞吐量。系统在车辆定位识别上达到84.3%准确率,在转向方向判断上达到76.4%准确率,优于传统模型。

原文摘要 · Abstract (English)

A robust and efficient traffic monitoring system is essential for smart cities and Intelligent Transportation Systems (ITS), using sensors and cameras to track vehicle movements, optimize traffic flow, reduce congestion, enhance road safety, and enable real-time adaptive traffic control. Traffic monitoring models must comprehensively understand dynamic urban conditions and provide an intuitive user interface for effective management. This research leverages the LLaVA visual grounding multimodal large language model (LLM) for traffic monitoring tasks on the real-time Quanser Interactive Lab simulation platform, covering scenarios like intersections, congestion, and collisions. Cameras placed at multiple urban locations collect real-time images from the simulation, which are fed into the LLaVA model with queries for analysis. An instance segmentation model integrated into the cameras highlights key elements such as vehicles and pedestrians, enhancing training and throughput. The system achieves 84.3% accuracy in recognizing vehicle locations and 76.4% in determining steering direction, outperforming traditional models.

交通监控多模态大模型实例分割

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。