用微调的QwenVL2模型+集成策略,拿下视频问答竞赛第一名
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
- 基于QwenVL2(7B)模型微调,结合测试时增强提升理解能力
- 在排行榜上达到0.7647的Top-1准确率,领先于其他方案
- 适合关注多模态视频理解与竞赛优化策略的研究者
本文介绍我们在第二届感知测试挑战赛多选题视频问答赛道中获得第一名的解决方案。该比赛提出了复杂的视频理解任务,要求模型能够准确理解并回答视频内容相关的问题。为应对这一挑战,我们采用强大的QwenVL2(7B)模型,并在其提供的训练集上进行微调。此外,我们还采用了模型集成策略和测试时增强技术以提升性能。通过持续优化,我们的方法在排行榜上达到了0.7647的Top-1准确率。
原文摘要 · Abstract (English)
In this report, we present our first-place solution to the Multiple-choice Video Question Answering (QA) track of The Second Perception Test Challenge. This competition posed a complex video understanding task, requiring models to accurately comprehend and answer questions about video content. To address this challenge, we leveraged the powerful QwenVL2 (7B) model and fine-tune it on the provided training set. Additionally, we employed model ensemble strategies and Test Time Augmentation to boost performance. Through continuous optimization, our approach achieved a Top-1 Accuracy of 0.7647 on the leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。