arXiv:2601.13719cs.CVcs.AI2026-01中稿 · CVPR被引 4

通过多模态实体一致性与智能搜索,实现长视频的连贯理解

Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search

论文配图:Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
图 1 · 摘自论文原文
  • 构建跨视听流的实体表示,分层组织视频内容
  • 动态检索与推理使长视频理解准确率达84.1%(LVBench)
  • 适合需要细粒度追踪与全局叙事的任务场景

长视频理解对视觉-语言模型构成严峻挑战,主要源于极长上下文窗口。现有基于简单分块与检索增强生成的方法通常导致信息碎片化和全局连贯性丢失。本文提出HAVEN框架,通过融合视听实体一致性与分层视频索引结合智能搜索机制,实现连贯、全面的推理。首先,通过跨视觉与听觉流的实体级表征保持语义一致性,并将内容组织为涵盖全局摘要、场景、片段和实体层级的结构化层次;其次,采用智能搜索机制在各层级间动态检索与推理,支持连贯的叙事重建与细粒度实体追踪。大量实验表明,该方法在时间连贯性、实体一致性和检索效率方面表现优异,在LVBench上达到84.1%的整体准确率,尤其在挑战性的推理类别中达80.1%。结果验证了结构化多模态推理在长视频综合、上下文一致理解中的有效性。

原文摘要 · Abstract (English)

Long video understanding presents significant challenges for vision-language models due to extremely long context windows. Existing solutions relying on naive chunking strategies with retrieval-augmented generation, typically suffer from information fragmentation and a loss of global coherence. We present HAVEN, a unified framework for long-video understanding that enables coherent and comprehensive reasoning by integrating audiovisual entity cohesion and hierarchical video indexing with agentic search. First, we preserve semantic consistency by integrating entity-level representations across visual and auditory streams, while organizing content into a structured hierarchy spanning global summary, scene, segment, and entity levels. Then we employ an agentic search mechanism to enable dynamic retrieval and reasoning across these layers, facilitating coherent narrative reconstruction and fine-grained entity tracking. Extensive experiments demonstrate that our method achieves good temporal coherence, entity consistency, and retrieval efficiency, establishing a new state-of-the-art with an overall accuracy of 84.1% on LVBench. Notably, it achieves outstanding performance in the challenging reasoning category, reaching 80.1%. These results highlight the effectiveness of structured, multimodal reasoning for comprehensive and context-consistent understanding of long-form videos.

长视频理解多模态智能搜索实体追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。