arXiv:2603.17840cs.CV2026-03被引 4

梳理视频理解三大方向,揭示从分治到统一建模的演进趋势。

Video Understanding: From Geometry and Semantics to Unified Models

  • 按几何、语义与统一模型三维度组织视频理解研究
  • 指出从单一任务模型转向可适配多目标的统一范式
  • 适合关注视频基础模型发展的研究人员参考

视频理解旨在使模型能够感知、推理并互动于动态视觉世界。相较于图像理解,视频理解需建模时间动态与演变的视觉上下文,对时空推理要求更高,是计算机视觉的核心问题。本文综述将文献分为三个互补视角:低层视频几何理解、高层语义理解以及统一视频理解模型。进一步指出研究正从孤立的任务专用流程向可适应多种下游目标的统一建模范式转变,为近期进展提供系统性视图。通过整合这些视角,本综述构建了视频理解演进的清晰地图,总结关键建模趋势与设计原则,并展望构建鲁棒、可扩展、统一的视频基础模型所面临的开放挑战。

原文摘要 · Abstract (English)

Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently requires modeling temporal dynamics and evolving visual context, placing stronger demands on spatiotemporal reasoning and making it a foundational problem in computer vision. In this survey, we present a structured overview of video understanding by organizing the literature into three complementary perspectives: low-level video geometry understanding, high-level semantic understanding, and unified video understanding models. We further highlight a broader shift from isolated, task-specific pipelines toward unified modeling paradigms that can be adapted to diverse downstream objectives, enabling a more systematic view of recent progress. By consolidating these perspectives, this survey provides a coherent map of the evolving video understanding landscape, summarizes key modeling trends and design principles, and outlines open challenges toward building robust, scalable, and unified video foundation models.

视频理解统一模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。