分任务训练双模型,提升交通视频分析的准确与全面。
Task-Specific Dual-Model Framework for Comprehensive Traffic Safety Video Description and Analysis
- 用两个模型分别专注描述与问答,减少干扰。
- 在WTS数据集上达45.76的S2分数,排名第10。
- 分离训练比联合训练提升8.6%问答准确率。
交通安全管理需要复杂的视频理解来捕捉细微行为模式并生成全面描述以预防事故。本文提出一种独特的双模型框架,通过任务特异性优化,充分利用VideoLLaMA与Qwen2.5-VL的互补优势。核心思路是将字幕生成与视觉问答(VQA)任务的训练分离,以减少任务干扰,使各模型更专注。实验表明,VideoLLaMA在时间推理上表现优异,取得1.1001的CIDEr得分;Qwen2.5-VL在视觉理解上更优,达到60.80%的VQA准确率。在WTS数据集上的大量实验显示,该方法在2025 AI City Challenge Track 2中获得45.7572的S2分数,位列第10。消融实验证明,分离训练策略在保持字幕质量的同时,使VQA准确率比联合训练高出8.6%。
原文摘要 · Abstract (English)
Traffic safety analysis requires complex video understanding to capture fine-grained behavioral patterns and generate comprehensive descriptions for accident prevention. In this work, we present a unique dual-model framework that strategically utilizes the complementary strengths of VideoLLaMA and Qwen2.5-VL through task-specific optimization to address this issue. The core insight behind our approach is that separating training for captioning and visual question answering (VQA) tasks minimizes task interference and allows each model to specialize more effectively. Experimental results demonstrate that VideoLLaMA is particularly effective in temporal reasoning, achieving a CIDEr score of 1.1001, while Qwen2.5-VL excels in visual understanding with a VQA accuracy of 60.80\%. Through extensive experiments on the WTS dataset, our method achieves an S2 score of 45.7572 in the 2025 AI City Challenge Track 2, placing 10th on the challenge leaderboard. Ablation studies validate that our separate training strategy outperforms joint training by 8.6\% in VQA accuracy while maintaining captioning quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。