arXiv:2510.09266cs.CL2025-10被引 1

构建细粒度视频多模态检索生成基准,揭示模型捕捉细微信息能力不足

CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation

  • 构建5360个开放问答对,覆盖图表报告、新闻播报等高密度多模态场景
  • 发现主流模型在长时序视频中难以捕捉瞬时关键细节,性能受限于细粒度理解
  • 提出自适应视觉增强框架,动态提升采样密度并智能调用外部工具

多模态检索增强生成(MRAG)使多模态大语言模型(MLLMs)能利用外部多模态证据生成回答,已有众多基于视频的MRAG基准用于评估模型在检索与生成阶段的能力。然而,现有基准在模态覆盖和格式多样性上仍显不足,常局限于单模态或粗粒度场景理解。为此,我们提出CFVBench,一个大规模、人工验证的基准,基于599个公开视频构建,生成5,360个开放式问答对。该基准涵盖高密度格式与多领域,如图表密集报告、新闻广播和软件教程,要求模型在长时序跨度内检索并推理细粒度多模态信息。通过CFVBench,我们系统评估了7种检索方法和14种主流MLLMs,揭示关键瓶颈:当前模型(包括GPT-5或Gemini)难以捕捉短暂但关键的细粒度多模态细节。为缓解此问题,我们提出自适应视觉增强(AVR)框架,可动态增加帧采样密度并在必要时选择性调用外部工具。实验表明,AVR显著提升所有被测MLLMs在细粒度多模态理解上的表现。

原文摘要 · Abstract (English)

Multimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence, and numerous video-based MRAG benchmarks have been proposed to evaluate model capabilities across retrieval and generation stages. However, existing benchmarks remain limited in modality coverage and format diversity, often focusing on single- or limited-modality tasks, or coarse-grained scene understanding. To address these gaps, we introduce CFVBench, a large-scale, manually verified benchmark constructed from 599 publicly available videos, yielding 5,360 open-ended QA pairs. CFVBench spans high-density formats and domains such as chart-heavy reports, news broadcasts, and software tutorials, requiring models to retrieve and reason over long temporal video spans while maintaining fine-grained multimodal information. Using CFVBench, we systematically evaluate 7 retrieval methods and 14 widely-used MLLMs, revealing a critical bottleneck: current models (even GPT5 or Gemini) struggle to capture transient yet essential fine-grained multimodal details. To mitigate this, we propose Adaptive Visual Refinement (AVR), a simple yet effective framework that adaptively increases frame sampling density and selectively invokes external tools when necessary. Experiments show that AVR consistently enhances fine-grained multimodal comprehension and improves performance across all evaluated MLLMs

多模态视频理解生成评测检索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。