用多个模型协作解视频问答难题,效果远超单模型。
Team of One: Cracking Complex Video QA with Model Synergy
- 让不同视频语言模型按思维链分工协作,提升推理能力。
- 在CVRR-ES数据集上各项指标显著优于现有方法。
- 无需重训练,适合想快速提升视频理解能力的研究者。
我们提出一种新型开放性视频问答框架,旨在提升复杂现实场景下的推理深度与鲁棒性,基于CVRR-ES数据集进行评估。现有视频大模型(Video-LMMs)普遍存在上下文理解有限、时间建模薄弱、对模糊或组合式问题泛化能力差等问题。为此,我们设计了一种提示与响应融合机制,通过结构化思维链协调多个异构视频-语言模型(VLMs),每个模型针对不同推理路径进行优化。外部大型语言模型(LLM)作为评估与整合器,选择并融合最可靠的输出。大量实验表明,该方法在所有评价指标上均显著优于现有基线,展现出更强的泛化能力和鲁棒性。本方案具有轻量、可扩展优势,无需模型重训练,为未来视频大模型发展奠定坚实基础。
原文摘要 · Abstract (English)
We propose a novel framework for open-ended video question answering that enhances reasoning depth and robustness in complex real-world scenarios, as benchmarked on the CVRR-ES dataset. Existing Video-Large Multimodal Models (Video-LMMs) often exhibit limited contextual understanding, weak temporal modeling, and poor generalization to ambiguous or compositional queries. To address these challenges, we introduce a prompting-and-response integration mechanism that coordinates multiple heterogeneous Video-Language Models (VLMs) via structured chains of thought, each tailored to distinct reasoning pathways. An external Large Language Model (LLM) serves as an evaluator and integrator, selecting and fusing the most reliable responses. Extensive experiments demonstrate that our method significantly outperforms existing baselines across all evaluation metrics, showcasing superior generalization and robustness. Our approach offers a lightweight, extensible strategy for advancing multimodal reasoning without requiring model retraining, setting a strong foundation for future Video-LMM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。