arXiv:2512.12360cs.CVcs.CL2025-12中稿 · CVPR被引 18

用动态推理与分层记忆提升长视频理解效率

VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding

  • 设计自适应推理循环,边看边思考边记,减少冗余处理
  • 在多个基准上超越当前最优方法,且令牌消耗大幅降低
  • 适合需要高效处理长视频的智能系统开发者

长视频理解因时间跨度长、多模态线索密集而具有挑战性。尽管已有进展,多数方法仍依赖人工设计的推理流程或耗资源的视频预处理来引导多模态大模型自主推理。为克服这些局限,我们提出 VideoARM,一种基于分层记忆的代理式推理范式,用于长视频理解。不同于静态、全面的预处理,VideoARM 实现自适应、实时的代理推理与记忆构建。具体而言,它持续执行观察、思考、行动和记忆的循环,由控制器自主调用工具以粗到细的方式解析视频,显著降低令牌消耗。同时,分层多模态记忆持续捕捉并更新各层次线索,为控制器提供精准上下文支持。在主流基准上的实验表明,VideoARM 在性能上超越当前最优方法 DVD,同时大幅减少长视频的令牌消耗。

原文摘要 · Abstract (English)

Long-form video understanding remains challenging due to the extended temporal structure and dense multimodal cues. Despite recent progress, many existing approaches still rely on hand-crafted reasoning pipelines or employ token-consuming video preprocessing to guide MLLMs in autonomous reasoning. To overcome these limitations, we introduce VideoARM, an Agentic Reasoning-over-hierarchical-Memory paradigm for long-form video understanding. Instead of static, exhaustive preprocessing, VideoARM performs adaptive, on-the-fly agentic reasoning and memory construction. Specifically, VideoARM performs an adaptive and continuous loop of observing, thinking, acting, and memorizing, where a controller autonomously invokes tools to interpret the video in a coarse-to-fine manner, thereby substantially reducing token consumption. In parallel, a hierarchical multimodal memory continuously captures and updates multi-level clues throughout the operation of the agent, providing precise contextual information to support the controller in decision-making. Experiments on prevalent benchmarks demonstrate that VideoARM outperforms the state-of-the-art method, DVD, while significantly reducing token consumption for long-form videos.

视频理解代理推理长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。