arXiv:2508.13439cs.CVcs.AI2025-08

用多智能体协作生成交通视频标注,让小模型也能实时识风险。

Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference

  • 用大模型协同生成带思维链的多视角标注
  • 3B小模型在多项指标上逼近大模型表现
  • 适合边缘部署,实现实时交通风险监测

全面理解高速公路场景并可靠推断交通风险,对智能交通系统和自动驾驶至关重要。传统方法在真实复杂动态环境下常面临可扩展性和泛化能力不足的问题。为此,我们提出一种新型结构化提示与多智能体知识蒸馏框架,实现高质量交通场景标注和上下文风险评估的自动生成。该框架协调两个大视觉语言模型(GPT-4o 和 o3-mini),采用结构化思维链策略生成多角度输出,作为知识增强的伪标注,用于监督微调一个更小的学生模型。最终得到的轻量级 3B 规模模型 VISTA(Vision for Intelligent Scene and Traffic Analysis)能理解低分辨率交通视频,并生成语义准确、具备风险意识的描述。尽管参数量大幅减少,VISTA 在标准图文生成指标(BLEU-4、METEOR、ROUGE-L、CIDEr)上仍表现出色,证明有效知识蒸馏与多智能体监督可使轻量级模型掌握复杂推理能力。VISTA 的紧凑架构支持在边缘设备高效部署,实现无需大规模基础设施升级的实时风险监控。

原文摘要 · Abstract (English)

Comprehensive highway scene understanding and robust traffic risk inference are vital for advancing Intelligent Transportation Systems (ITS) and autonomous driving. Traditional approaches often struggle with scalability and generalization, particularly under the complex and dynamic conditions of real-world environments. To address these challenges, we introduce a novel structured prompting and knowledge distillation framework that enables automatic generation of high-quality traffic scene annotations and contextual risk assessments. Our framework orchestrates two large Vision-Language Models (VLMs): GPT-4o and o3-mini, using a structured Chain-of-Thought (CoT) strategy to produce rich, multi-perspective outputs. These outputs serve as knowledge-enriched pseudo-annotations for supervised fine-tuning of a much smaller student VLM. The resulting compact 3B-scale model, named VISTA (Vision for Intelligent Scene and Traffic Analysis), is capable of understanding low-resolution traffic videos and generating semantically faithful, risk-aware captions. Despite its significantly reduced parameter count, VISTA achieves strong performance across established captioning metrics (BLEU-4, METEOR, ROUGE-L, and CIDEr) when benchmarked against its teacher models. This demonstrates that effective knowledge distillation and structured multi-agent supervision can empower lightweight VLMs to capture complex reasoning capabilities. The compact architecture of VISTA facilitates efficient deployment on edge devices, enabling real-time risk monitoring without requiring extensive infrastructure upgrades.

交通视频知识蒸馏边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。