arXiv:2606.12125cs.CV2026-06被引 1

让长视频理解更高效,聚焦关键帧,保留时间连贯性

Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding

论文配图:Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding
图 1 · 摘自论文原文
  • 按查询选择重点片段,分作高保真焦点帧与上下文布局
  • 在不增加计算量前提下,最长视频任务提升9.1个百分点
  • 无需训练,适配多种视频大模型,特别适合长视频场景

长视频理解对多模态大语言模型仍是难题,因视频动辄数千帧,全量处理成本高昂。现有方法虽在有限视觉预算下构造紧凑输入,但多沿用帧为中心的范式,对所有保留内容采用相似表示,难以兼顾高保真视觉证据与广范围时间覆盖。为此,我们提出Q-Fold——一种无需训练的长视频输入构建框架。不以孤立帧为建模单元,而是基于连续时间片段,结合查询引导构建异构的焦点-上下文表征:相关片段作为高保真焦点帧保留,无关片段则折叠为保持时序结构的上下文布局。该方法既保留关键视觉信息,又实现广泛的时间覆盖,并更好维持短片段内局部时序连续性。在四个长视频基准上使用多个Video-MLLM进行实验,结果表明Q-Fold在不增加输入预算的前提下持续提升性能,尤其在超长视频任务中最高取得9.1个百分点的增益。

原文摘要 · Abstract (English)

Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames and are therefore expensive to process exhaustively. Existing methods usually construct compact visual inputs from long videos under a limited visual budget. However, most of them still follow a frame-centric paradigm and apply similar representations to retained content regardless of its importance. This makes it difficult to preserve both high-fidelity visual evidence and broad temporal coverage. To address this issue, we propose Q-Fold, a training-free input construction framework for long-video understanding. Instead of treating isolated frames as the basic modeling unit, Q-Fold operates on contiguous temporal segments and constructs a heterogeneous Focus--Context representation under query guidance. Query-relevant segments are preserved as high-fidelity Focus Frames, while less relevant segments are folded into chronology-preserving contextual layouts. In this way, Q-Fold preserves critical visual evidence and broad temporal coverage, while better maintaining local temporal continuity within short segments. Experiments on four long-video benchmarks with multiple Video-MLLMs show that Q-Fold consistently improves performance without increasing the input budget. Notably, it achieves gains of up to 9.1 percentage points on an ultra-long video benchmark. Code will be made publicly available.

长视频理解视觉推理多模态模型注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。