arXiv:2607.28516cs.CV2026-07

通过生成式隐变量聚合,让长视频理解更懂跨帧信息融合。

Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding

论文配图:Beyond Frame Selection: Generative Latent Evidence Aggregation for Long-Video Understanding
图 1 · 摘自论文原文
  • 选帧后用查询引导的隐变量分布整合跨帧证据
  • 8帧下提升4个基准平均分5.2点,最高增10.1点
  • 仅增加微量视频令牌开销,适合长视频任务

长视频理解通常将视频压缩为少量帧或视觉标记进行回答生成。现有紧凑流程注重保留相关视觉内容作为显式证据,但提供证据并不保证跨时刻互补线索被整合。本文核心思想是在生成前,将选定帧组织成与查询相关的跨帧证据。我们提出一个后选帧阶段的隐变量证据接口,实现为GenEvA(生成式隐变量证据聚合)框架。GenEvA利用查询条件化的证据分布,聚焦于相关帧,从帧级信息中形成紧凑的跨帧隐变量证据。由于跨帧整合并非总必要,该分布同时决定是否插入此隐变量补充。在四个基准和两个Video-MLLM模型上,GenEvA均持续优于匹配帧基线。在8帧设置下,其使四个基准的LLaVA-Video平均分提升5.2点,Qwen2.5-VL在LVBench上的准确率提升10.1点。这些提升仅需0.11%–0.40%的平均视频标记开销;分析进一步显示任务感知分配与自适应证据调用的优势。

原文摘要 · Abstract (English)

Long-video understanding commonly compresses videos into a small set of frames or visual tokens for answer generation. Existing compact pipelines focus on retaining relevant visual content as explicit evidence. Yet making evidence available does not ensure that complementary cues across moments are integrated for answering. Our key idea is to organize selected frames into query-relevant cross-frame evidence before generation. We formulate this post-selection stage as a latent evidence interface and instantiate it with GenEvA ($\textbf{Gen}erative$ $Latent$ $\textbf{Ev}idence$ $\textbf{A}ggregation$), a distribution-guided latent evidence aggregation framework. Specifically, GenEvA uses a query-conditioned evidence distribution to focus aggregation on relevant frames, forming compact cross-frame latent evidence from their frame-specific information. Since cross-frame integration is not always needed, the same distribution determines whether to insert this latent complement. Across four benchmarks and two Video-MLLM backbones, GenEvA consistently improves matched-frame baselines. At 8 frames, it raises the four-benchmark LLaVA-Video average by $+5.2$ points and Qwen2.5-VL accuracy on LVBench by $+10.1$ points. These gains require only $0.11\%$--$0.40\%$ average video-token overhead; analyses further show task-aware allocation and benefits from Adaptive Evidence Invocation.

长视频理解生成式聚合跨帧融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。