arXiv:2604.01002cs.CVcs.AI2026-04被引 1

基于查询的证据关键帧采样,提升长视频理解效率

Query-Conditioned Evidential Keyframe Sampling for MLLM-Based Long-Form Video Understanding

  • 以信息瓶颈理论为指导,最大化帧与问题间的条件互信息
  • 在严格令牌预算下,准确率超越现有方法,训练效率显著提升
  • 适合需要高效处理长视频的多模态大模型应用

多模态大语言模型在视频问答任务中表现优异,但受限于上下文长度和计算成本,难以直接应用于长视频。关键帧采样成为必要手段。现有方法通常依赖语义相关性或强化学习,要么忽略证据线索,要么存在组合优化效率低的问题。本文提出一种基于证据驱动的关键帧采样框架,建立在信息瓶颈理论基础上,将关键帧选择建模为最大化所选帧与问题之间的条件互信息,从而提供反映每帧对回答贡献的合理目标。为使该目标可计算,我们利用其结构特性,推导出可分解的优化形式,将子集选择简化为独立的帧级评分。进一步设计了一个查询条件的证据评分网络,通过对比学习目标高效估计证据重要性。在多个长视频理解基准上的实验表明,本方法在严格令牌预算下持续优于现有采样策略,同时显著提升训练效率。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have shown strong performance on video question answering, but their application to long-form videos is constrained by limited context length and computational cost, making keyframe sampling essential. Existing approaches typically rely on semantic relevance or reinforcement learning, which either fail to capture evidential clues or suffer from inefficient combinatorial optimization. In this work, we propose an evidence-driven keyframe sampling framework grounded in information bottleneck theory. We formulate keyframe selection as maximizing the conditional mutual information between selected frames and the query, providing a principled objective that reflects each frame's contribution to answering the question. To make this objective tractable, we exploit its structure to derive a decomposed optimization that reduces subset selection to independent frame-level scoring. We further introduce a query-conditioned evidence scoring network trained with a contrastive objective to estimate evidential importance efficiently. Experiments on long-form video understanding benchmarks show that our method consistently outperforms prior sampling strategies under strict token budgets, while significantly improving training efficiency.

视频理解关键帧采样多模态模型信息瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。