提出SlowFocus机制,让视频大模型更精准理解时间细节。
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
- 根据问题定位关键时间段,密集采样提取高频局部特征
- 多频率混合注意力融合局部细节与全局上下文,提升时间理解能力
- 专设细粒度视频理解基准FineAction-CGR,适合视频分析研究者
大型语言模型在文本理解方面表现出色,推动其向视频大模型(Vid-LLMs)扩展以分析视频数据。然而,现有Vid-LLMs难以同时保持高质量的帧级语义信息(即每帧足够多的视觉标记)和全面的视频级时间信息(即每视频足够多的采样帧),限制了其向细粒度视频理解的发展。为此,我们提出SlowFocus机制,显著提高等效采样频率而不损失帧级视觉标记质量。该机制首先根据问题定位相关时间片段,然后在此片段上进行密集采样以提取局部高频特征。进一步引入多频率混合注意力模块,将这些局部高频细节与全局低频上下文聚合,增强时间理解能力。此外,为适配此新机制,我们设计了一套训练策略,以强化时间定位与精细时间推理能力。同时,我们构建了FineAction-CGR基准,专门用于评估Vid-LLMs处理细粒度时间理解任务的能力。大量实验表明,该机制在现有公开视频理解基准及自建的FineAction-CGR上均表现优越。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated exceptional capabilities in text understanding, which has paved the way for their expansion into video LLMs (Vid-LLMs) to analyze video data. However, current Vid-LLMs struggle to simultaneously retain high-quality frame-level semantic information (i.e., a sufficient number of tokens per frame) and comprehensive video-level temporal information (i.e., an adequate number of sampled frames per video). This limitation hinders the advancement of Vid-LLMs towards fine-grained video understanding. To address this issue, we introduce the SlowFocus mechanism, which significantly enhances the equivalent sampling frequency without compromising the quality of frame-level visual tokens. SlowFocus begins by identifying the query-related temporal segment based on the posed question, then performs dense sampling on this segment to extract local high-frequency features. A multi-frequency mixing attention module is further leveraged to aggregate these local high-frequency details with global low-frequency contexts for enhanced temporal comprehension. Additionally, to tailor Vid-LLMs to this innovative mechanism, we introduce a set of training strategies aimed at bolstering both temporal grounding and detailed temporal reasoning capabilities. Furthermore, we establish FineAction-CGR, a benchmark specifically devised to assess the ability of Vid-LLMs to process fine-grained temporal understanding tasks. Comprehensive experiments demonstrate the superiority of our mechanism across both existing public video understanding benchmarks and our proposed FineAction-CGR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。