arXiv:2507.09068cs.CVcs.AI2025-07被引 2

提出无限视频理解新方向,突破长视频处理极限。

Infinite Video Understanding

  • 以持续处理任意时长视频为目标,探索新型架构与表示方法。
  • 解决超长视频中的计算瓶颈、时序连贯性与细节保持难题。
  • 适合关注长视频理解、持续学习与智能代理的 researchers。

大语言模型及其多模态扩展在视频理解领域取得显著进展,但面对分钟级乃至小时级以上的视频内容仍存在根本挑战。尽管近期工作如 Video-XL-2 提出高效架构,霍普位置编码(HoPE)和 VideoRoPE++ 等改进时空建模,当前最先进模型在处理超长序列产生的海量视觉标记时,仍面临严重的计算与内存限制。同时,维持时间连贯性、追踪复杂事件及长期保留细粒度信息依然是重大难题,即使在 Deep Video Discovery 等智能推理系统中也未完全解决。本文主张,将「无限视频理解」——即模型对任意长度、可能永不停止的视频数据实现持续处理、理解与推理——作为多媒体研究的下一前沿目标。这一愿景为多媒体及更广泛的人工智能研究提供关键指引,推动流式架构、持久记忆机制、分层自适应表征、以事件为中心的推理以及新型评估范式的发展。本文借鉴长视频/超长视频理解及相关领域成果,系统梳理核心挑战与关键研究方向。

原文摘要 · Abstract (English)

The rapid advancements in Large Language Models (LLMs) and their multimodal extensions (MLLMs) have ushered in remarkable progress in video understanding. However, a fundamental challenge persists: effectively processing and comprehending video content that extends beyond minutes or hours. While recent efforts like Video-XL-2 have demonstrated novel architectural solutions for extreme efficiency, and advancements in positional encoding such as HoPE and VideoRoPE++ aim to improve spatio-temporal understanding over extensive contexts, current state-of-the-art models still encounter significant computational and memory constraints when faced with the sheer volume of visual tokens from lengthy sequences. Furthermore, maintaining temporal coherence, tracking complex events, and preserving fine-grained details over extended periods remain formidable hurdles, despite progress in agentic reasoning systems like Deep Video Discovery. This position paper posits that a logical, albeit ambitious, next frontier for multimedia research is Infinite Video Understanding -- the capability for models to continuously process, understand, and reason about video data of arbitrary, potentially never-ending duration. We argue that framing Infinite Video Understanding as a blue-sky research objective provides a vital north star for the multimedia, and the wider AI, research communities, driving innovation in areas such as streaming architectures, persistent memory mechanisms, hierarchical and adaptive representations, event-centric reasoning, and novel evaluation paradigms. Drawing inspiration from recent work on long/ultra-long video understanding and several closely related fields, we outline the core challenges and key research directions towards achieving this transformative capability.

视频理解长视频AI前沿持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。