arXiv:2510.10682cs.CV2025-10

通过动态建模与跨时交互,实现在线动作理解的检测与预测统一优化。

Action-Dynamics Modeling and Cross-Temporal Interaction for Online Action Understanding

  • 用关键状态压缩帧序列,减少冗余信息。
  • 构建多维状态转移图,捕捉复杂场景中的动作动态与意图线索。
  • 跨时交互机制融合过去、当前与未来信息,适合医疗行为分析等实时场景。

动作理解(包含动作检测与预测)在众多实际应用中至关重要。然而,未剪辑视频常含有大量冗余信息和噪声,且现有方法常忽略主体意图对动作的影响。为此,本文提出状态特定模型(SSM),统一优化动作检测与预测任务。该框架通过关键状态记忆压缩模块将帧序列压缩为关键状态,降低信息冗余;通过动作模式学习模块构建具有多维边的状态转移图,建模复杂场景下的动作动态,并生成潜在未来线索以表征意图;此外,跨时交互模块通过跨时间的相互作用,建模意图与过去、当前信息的相互影响,从而优化当前与未来特征,实现检测与预测的同步。在多个基准数据集(包括EPIC-Kitchens-100、THUMOS'14、TVSeries及自建帕金森病小鼠行为数据集PDMB)上的实验表明,所提框架性能优于现有最先进方法。结果凸显动作动态学习与跨时交互的重要性,为后续动作理解研究奠定基础。

原文摘要 · Abstract (English)

Action understanding, encompassing action detection and anticipation, plays a crucial role in numerous practical applications. However, untrimmed videos are often characterized by substantial redundant information and noise. Moreover, in modeling action understanding, the influence of the agent's intention on the action is often overlooked. Motivated by these issues, we propose a novel framework called the State-Specific Model (SSM), designed to unify and enhance both action detection and anticipation tasks. In the proposed framework, the Critical State-Based Memory Compression module compresses frame sequences into critical states, reducing information redundancy. The Action Pattern Learning module constructs a state-transition graph with multi-dimensional edges to model action dynamics in complex scenarios, on the basis of which potential future cues can be generated to represent intention. Furthermore, our Cross-Temporal Interaction module models the mutual influence between intentions and past as well as current information through cross-temporal interactions, thereby refining present and future features and ultimately realizing simultaneous action detection and anticipation. Extensive experiments on multiple benchmark datasets -- including EPIC-Kitchens-100, THUMOS'14, TVSeries, and the introduced Parkinson's Disease Mouse Behaviour (PDMB) dataset -- demonstrate the superior performance of our proposed framework compared to other state-of-the-art approaches. These results highlight the importance of action dynamics learning and cross-temporal interactions, laying a foundation for future action understanding research.

动作理解动态建模意图识别实时分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。