arXiv:2503.08576cs.CV2025-03被引 7

用问答相关帧采样提升长视频理解评测准确性

RAG-Adapter: A Plug-and-Play RAG-enhanced Framework for Long Video Understanding

  • 根据问题动态选择关键帧,减少信息丢失
  • 在Video-MME上使GPT-4o准确率提升9.3%
  • 适合作为长视频评测的通用增强方案

多模态大语言模型在视频理解方面进展迅速。为有效评估其能力,提出了如Video-MME和MLVU等长视频理解基准。然而,这些基准直接采用均匀帧采样进行测试,导致显著信息丢失,影响评估结果对MLLM真实能力的反映。为此,我们提出RAG-Adapter,一种即插即用框架,通过采样与问题最相关的帧来减少测试过程中的信息损失。此外,我们构建了MMAT数据集,并引入分组监督对比学习(GCL)方法,通过微调进一步提升RAG-Adapter的采样效果。我们在多个基准上测试了多种基线MLLM,发现RAG-Adapter采样在所有情况下均优于均匀采样(例如,GPT-4o在Video-MME上的准确率提升9.3%),提供了一种更准确的长视频评测方法。

原文摘要 · Abstract (English)

Multi-modal Large Language Models (MLLMs) capable of video understanding are advancing rapidly. To effectively assess their video comprehension capabilities, long video understanding benchmarks, such as Video-MME and MLVU, are proposed. However, these benchmarks directly use uniform frame sampling for testing, which results in significant information loss and affects the accuracy of the evaluations in reflecting the true abilities of MLLMs. To address this, we propose RAG-Adapter, a plug-and-play framework that reduces information loss during testing by sampling frames most relevant to the given question. Additionally, we introduce a Grouped-supervised Contrastive Learning (GCL) method to further enhance sampling effectiveness of RAG-Adapter through fine-tuning on our constructed MMAT dataset. Finally, we test numerous baseline MLLMs on various video understanding benchmarks, finding that RAG-Adapter sampling consistently outperforms uniform sampling (e.g., Accuracy of GPT-4o increases by 9.3 percent on Video-MME), providing a more accurate testing method for long video benchmarks.

视频理解RAG评测优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。