用多专家架构提升足球问答准确率至95%
MSUE: Multi-Modal Soccer Understanding Expert

- 用视觉语言模型生成多样化足球问答数据
- 多专家动态分配,准确率达0.95,排名第三
- 适合关注视频理解与多模态问答的开发者
本文提出针对2026年SoccerNet VQA挑战赛的解决方案。首先,构建一种基于视觉-语言模型(VLM)的低成本数据合成管道,将原始领域数据系统性重构为多样化的问答样本,涵盖简短答案与长文本回答。其次,提出MSUE多专家问答架构,利用大语言模型(LLM)动态将问题分发至文本、图像和视频专家。这三个专家分别由强文本基线Gemini3-Flash、微调后的Qwen3-VL以及外部知识库实现,协同提升问答性能。MSUE在挑战赛基准上取得0.95的准确率,位列排行榜第三。
原文摘要 · Abstract (English)
This paper presents our solution to the 2026 SoccerNet VQA Challenge. We first develop a cost-effective data synthesis pipeline driven by a Vision-Language Model (VLM), which systematically restructures raw domain data into diverse VQA samples, including concise answers and long-form responses. Second, we propose MSUE, a multi-expert question answering architecture that employs a Large Language Model (LLM) to dynamically dispatch questions to text, image, and video experts. These experts are instantiated as a strong text baseline Gemini3-Flash, a fine-tuned Qwen3-VL, and an external knowledge base, respectively, working collaboratively to enhance VQA performance. MSUE achieves an accuracy of \textbf{0.95} on the challenge benchmark, securing third place in the leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。