提出双路径网络,更好识别1-3秒的细微动作。
Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition

- 双路并行处理:先空间后时间,或先时间后空间,适配不同动作特征。
- 在MA-52和iMiGUE数据集上达领先性能,尤其iMiGUE为当前最佳。
- 每部分自适应选择最优路径,适合细粒度视频理解研究者。
微动作是持续1-3秒的细微局部动作,如抓头或敲指,对社交交流至关重要,广泛存在于自然交互中,是细粒度视频理解的关键,但现有视觉系统对其理解不足。我们发现根本挑战在于:微动作具有多样化的时空特性,有的由空间结构定义,有的依赖时间动态表现。现有方法仅采用单一时空分解,无法应对这种多样性。为此,我们提出双路径网络,通过并行的空间-时间(ST)与时间-空间(TS)路径处理解剖学基础的空间实体。ST路径先捕捉空间结构再建模时间动态,而TS路径则反向处理以强调时间动态。不采用固定融合,而是引入实体级自适应路由,使每个身体部位学习其最优处理偏好,并辅以相互动作一致性(MAC)损失,确保跨路径一致性。大量实验表明,该方法在MA-52数据集上表现优异,在iMiGUE数据集上达到当前最佳结果。本工作揭示,针对微动作内在复杂性进行架构适配,对推进细粒度视频理解至关重要。
原文摘要 · Abstract (English)
Micro-actions are subtle, localized movements lasting 1-3 seconds such as scratching one's head or tapping fingers. Such subtle actions are essential for social communication, ubiquitously used in natural interactions, and thus critical for fine-grained video understanding, yet remain poorly understood by current computer vision systems. We identify a fundamental challenge: micro-actions exhibit diverse spatio-temporal characteristics where some are defined by spatial configurations while others manifest through temporal dynamics. Existing methods that commit to a single spatio-temporal decomposition cannot accommodate this diversity. We propose a dual-path network that processes anatomically-grounded spatial entities through parallel Spatial-Temporal (ST) and Temporal-Spatial (TS) pathways. The ST path captures spatial configurations before modeling temporal dynamics, while the TS path inverts this order to prioritize temporal dynamics. Rather than fixed fusion, we introduce entity-level adaptive routing where each body part learns its optimal processing preference, complemented by Mutual Action Consistency (MAC) loss that enforces cross-path coherence. Extensive experiments demonstrate competitive performance on MA-52 dataset and state-of-the-art results on iMiGUE dataset. Our work reveals that architectural adaptation to the inherent complexity of micro-actions is essential for advancing fine-grained video understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。