arXiv:2411.15106cs.CVcs.AI2024-11中稿 · IJCV被引 6

系统梳理动作理解的进展与挑战,覆盖三类时序任务。

About Time: Advances, Challenges, and Outlooks of Action Understanding

  • 按时间跨度分为识别、预测、预报三类任务
  • 涵盖单模态与多模态动作理解的主流方法
  • 适合关注视频理解前沿的研究者与开发者

近年来视频动作理解取得显著进展。数据集规模扩大、多样性提升以及算力增长,推动性能跃升与任务多样化。当前系统可提供视频场景的粗粒度与细粒度描述,定位查询对应片段,补全未观测视频内容,并实现跨模态上下文预测。本文全面综述了单模态与多模态动作理解在各类任务中的进展。重点分析普遍挑战,梳理常用数据集,回顾关键文献并聚焦近期突破。根据时间范围,将任务划分为三类:(1) 完整观测动作的识别任务;(2) 部分观测动作的预测任务;(3) 未观测后续动作的预报任务。该划分有助于识别动作建模与视频表征的关键挑战。最后,提出未来发展方向以应对现有不足。

原文摘要 · Abstract (English)

We have witnessed impressive advances in video action understanding. Increased dataset sizes, variability, and computation availability have enabled leaps in performance and task diversification. Current systems can provide coarse- and fine-grained descriptions of video scenes, extract segments corresponding to queries, synthesize unobserved parts of videos, and predict context across multiple modalities. This survey comprehensively reviews advances in uni- and multi-modal action understanding across a range of tasks. We focus on prevalent challenges, overview widely adopted datasets, and survey seminal works with an emphasis on recent advances. We broadly distinguish between three temporal scopes: (1) recognition tasks of actions observed in full, (2) prediction tasks for ongoing partially observed actions, and (3) forecasting tasks for subsequent unobserved action(s). This division allows us to identify specific action modeling and video representation challenges. Finally, we outline future directions to address current shortcomings.

动作理解视频分析多模态综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。