arXiv:2507.22421cs.CVcs.AI2025-07被引 2

统一框架实现动作识别与目标追踪,实时高效且精度领先。

Efficient Spatial-Temporal Modeling for Real-Time Video Analysis: A Unified Framework for Action Recognition and Object Tracking

  • 引入分层注意力机制,动态聚焦时空序列中的关键区域。
  • 在UCF-101等数据集上,动作识别准确率提升3.2%,追踪精度提高2.8%。
  • 推理速度比现有方法快40%,适合嵌入式等资源受限场景。

实时视频分析在计算机视觉中仍具挑战性,需高效处理空间与时间信息同时保持计算效率。现有方法常难以平衡精度与速度,尤其在资源受限环境下。本文提出一种统一框架,结合先进时空建模技术,实现动作识别与目标追踪的联合优化。基于并行序列建模进展,引入新颖的分层注意力机制,自适应关注时序中相关空间区域。在UCF-101、HMDB-51和MOT17数据集上的实验表明,该方法在标准基准上达到当前最优性能,推理速度比现有方法快40%,动作识别准确率提升3.2%,追踪精度提高2.8%。

原文摘要 · Abstract (English)

Real-time video analysis remains a challenging problem in computer vision, requiring efficient processing of both spatial and temporal information while maintaining computational efficiency. Existing approaches often struggle to balance accuracy and speed, particularly in resource-constrained environments. In this work, we present a unified framework that leverages advanced spatial-temporal modeling techniques for simultaneous action recognition and object tracking. Our approach builds upon recent advances in parallel sequence modeling and introduces a novel hierarchical attention mechanism that adaptively focuses on relevant spatial regions across temporal sequences. We demonstrate that our method achieves state-of-the-art performance on standard benchmarks while maintaining real-time inference speeds. Extensive experiments on UCF-101, HMDB-51, and MOT17 datasets show improvements of 3.2% in action recognition accuracy and 2.8% in tracking precision compared to existing methods, with 40% faster inference time.

视频分析动作识别目标追踪实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。