HierarQ通过分层查询机制,实现长视频理解的高效精准分析。
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
- 分两路处理帧级与场景级信息,结合任务感知调制器增强理解。
- 在10个视频基准上超越主流模型,长视频任务表现更优。
- 适合需要细粒度视频分析的研究者和开发者使用。
尽管多模态大语言模型取得进展,当前方法在中长视频理解方面仍受限于帧数和上下文长度。现有模型常依赖帧采样,易遗漏关键信息且缺乏任务相关性。为此,我们提出HierarQ,一种任务感知的分层查询框架,通过序列化处理帧,避免帧采样并突破大模型上下文长度限制。引入轻量级双流语言引导特征调制器:实体流在短上下文中捕捉帧级目标信息,场景流则在更长时段内识别其交互关系。每一流均配备专用记忆库,使提出的分层查询变压器(HierarQ)能有效捕获短时与长时上下文。在10个涵盖视频理解、问答与字幕生成任务的基准上进行广泛评估,结果显示HierarQ在多数数据集上达到领先性能,证明其在综合视频分析中的鲁棒性与高效性。
原文摘要 · Abstract (English)
Despite advancements in multimodal large language models (MLLMs), current approaches struggle in medium-to-long video understanding due to frame and context length limitations. As a result, these models often depend on frame sampling, which risks missing key information over time and lacks task-specific relevance. To address these challenges, we introduce HierarQ, a task-aware hierarchical Q-Former based framework that sequentially processes frames to bypass the need for frame sampling, while avoiding LLM's context length limitations. We introduce a lightweight two-stream language-guided feature modulator to incorporate task awareness in video understanding, with the entity stream capturing frame-level object information within a short context and the scene stream identifying their broader interactions over longer period of time. Each stream is supported by dedicated memory banks which enables our proposed Hierachical Querying transformer (HierarQ) to effectively capture short and long-term context. Extensive evaluations on 10 video benchmarks across video understanding, question answering, and captioning tasks demonstrate HierarQ's state-of-the-art performance across most datasets, proving its robustness and efficiency for comprehensive video analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。