用多模态检索增强生成,提升视频问答中开放问题的知识推理能力
Open-Ended and Knowledge-Intensive Video Question Answering
- 采用多模态检索增强生成框架,融合视觉与外部知识
- 在KnowIT数据集上多选题准确率提升17.5%,达新纪录
- 适合需要跨模态知识推理的开放性视频问答任务
仅依赖视觉内容的视频问答已取得进展,但涉及外部知识的开放性问题仍是重大挑战。本文聚焦知识密集型视频问答(KI-VideoQA),通过多模态检索增强生成方法应对开放性问题,而非仅限于选择题。我们系统评估了前沿检索与视觉语言模型在零样本和微调配置下的表现,分析不同信息源与模态的交互、多模态上下文整合策略,以及查询构建与检索结果利用之间的动态关系。结果表明,检索增强虽有潜力,但性能高度依赖模态选择与检索方法。研究强调查询设计与检索深度优化对知识有效整合的关键作用。所提方法在KnowIT VQA数据集上实现多选题准确率17.5%的显著提升,刷新当前最优水平。
原文摘要 · Abstract (English)
Video question answering that requires external knowledge beyond the visual content remains a significant challenge in AI systems. While models can effectively answer questions based on direct visual observations, they often falter when faced with questions requiring broader contextual knowledge. To address this limitation, we investigate knowledge-intensive video question answering (KI-VideoQA) through the lens of multi-modal retrieval-augmented generation, with a particular focus on handling open-ended questions rather than just multiple-choice formats. Our comprehensive analysis examines various retrieval augmentation approaches using cutting-edge retrieval and vision language models, testing both zero-shot and fine-tuned configurations. We investigate several critical dimensions: the interplay between different information sources and modalities, strategies for integrating diverse multi-modal contexts, and the dynamics between query formulation and retrieval result utilization. Our findings reveal that while retrieval augmentation shows promise in improving model performance, its success is heavily dependent on the chosen modality and retrieval methodology. The study also highlights the critical role of query construction and retrieval depth optimization in effective knowledge integration. Through our proposed approach, we achieve a substantial 17.5% improvement in accuracy on multiple choice questions in the KnowIT VQA dataset, establishing new state-of-the-art performance levels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。