arXiv:2502.07277cs.CVcs.AI2025-02

探索深度神经网络在视频时空特征分析中的应用

Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis

  • 基于深度神经网络提取视频时空特征
  • 系统梳理主流视频理解模型与结构设计
  • 对比分析典型数据集,适合研究者参考

视频已成为在线信息传播的主要方式,推动了对视频内容分析算法的需求持续增长。这些算法需从视频中提取并分类特征,以描述其中的事件与物体。深度神经网络在特征提取与视频描述任务中展现出良好效果。本文探讨视频中的时空特征,综述深度神经网络在视频理解领域的最新进展,分析主流模型的结构设计、核心问题及解决方案,并对比重要视频理解与动作识别数据集。

原文摘要 · Abstract (English)

It's no secret that video has become the primary way we share information online. That's why there's been a surge in demand for algorithms that can analyze and understand video content. It's a trend going to continue as video continues to dominate the digital landscape. These algorithms will extract and classify related features from the video and will use them to describe the events and objects in the video. Deep neural networks have displayed encouraging outcomes in the realm of feature extraction and video description. This paper will explore the spatiotemporal features found in videos and recent advancements in deep neural networks in video understanding. We will review some of the main trends in video understanding models and their structural design, the main problems, and some offered solutions in this topic. We will also review and compare significant video understanding and action recognition datasets.

视频理解时空分析深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。