arXiv:2601.11359cs.CVcs.AI2026-01中稿 · ICASSP2026被引 4

通过慢快帧采样提升长视频理解,无需训练即可提速降耗

Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding

  • 多查询推理+片段级慢快采样,兼顾局部细节与全局上下文
  • 在多个数据集上最高提升6.9%准确率,推理耗时减少50%
  • 适用于长视频任务,尤其适合资源受限的多模态大模型部署

多模态大语言模型(MLLM)在视频理解方面取得显著进展,但在长视频上的表现仍受计算限制和帧选择不佳的影响。本文提出无需训练的Think-Clip-Sample(TCS)框架,包含两个核心组件:(i) 多查询推理,生成多个查询以捕捉问题与视频的互补信息;(ii) 片段级慢快采样,自适应平衡密集局部细节与稀疏全局上下文。在MLVU、LongVideoBench和VideoMME上的大量实验表明,TCS能持续提升不同MLLM的表现,准确率最高提升6.9%,且在保持相近精度的同时将推理时间降低50%,充分展现了其在长视频理解中的高效性与有效性。

原文摘要 · Abstract (English)

Recent progress in multi-modal large language models (MLLMs) has significantly advanced video understanding. However, their performance on long-form videos remains limited by computational constraints and suboptimal frame selection. We present Think-Clip-Sample (TCS), a training-free framework that enhances long video understanding through two key components: (i) Multi-Query Reasoning, which generates multiple queries to capture complementary aspects of the question and video; and (ii) Clip-level Slow-Fast Sampling, which adaptively balances dense local details and sparse global context. Extensive experiments on MLVU, LongVideoBench, and VideoMME demonstrate that TCS consistently improves performance across different MLLMs, boosting up to 6.9% accuracy, and is capable of achieving comparable accuracy with 50% fewer inference time cost, highlighting both efficiency and efficacy of TCS on long video understanding.

视频理解多模态推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。