通过定位前景优化提升视频重复动作计数的鲁棒性。
Localization-Aware Multi-Scale Representation Learning for Repetitive Action Counting
- 引入前景定位目标,增强时序特征对重复动作的感知能力。
- 在RepCountA和UCFRep数据集上实现更优计数精度,抗噪声干扰强。
- 适合需要高鲁棒性动作计数的应用场景,如工业质检、运动分析。
重复动作计数(RAC)旨在无需示例的情况下估计视频中类无关动作的发生次数。现有方法依赖帧间相似性表示进行周期预测,但易受动作中断、不一致等常见噪声影响,导致真实场景下性能不佳。本文提出一种定位感知多尺度表征学习框架(LMRL),通过引入重复动作前景定位(RFL)方法,粗粒度识别周期性动作并融合全局语义信息,增强表征能力;同时设计尺度自适应的多尺度周期感知表示(MPR),以适应不同动作频率,学习更灵活的时序相关性。两个模块联合优化,显著降低噪声影响,提升计数准确性。该框架具备良好可扩展性与内容适应性。在RepCountA和UCFRep数据集上的实验表明,所提方法有效解决了重复动作计数问题。
原文摘要 · Abstract (English)
Repetitive action counting (RAC) aims to estimate the number of class-agnostic action occurrences in a video without exemplars. Most current RAC methods rely on a raw frame-to-frame similarity representation for period prediction. However, this approach can be significantly disrupted by common noise such as action interruptions and inconsistencies, leading to sub-optimal counting performance in realistic scenarios. In this paper, we introduce a foreground localization optimization objective into similarity representation learning to obtain more robust and efficient video features. We propose a Localization-Aware Multi-Scale Representation Learning (LMRL) framework. Specifically, we apply a Multi-Scale Period-Aware Representation (MPR) with a scale-specific design to accommodate various action frequencies and learn more flexible temporal correlations. Furthermore, we introduce the Repetition Foreground Localization (RFL) method, which enhances the representation by coarsely identifying periodic actions and incorporating global semantic information. These two modules can be jointly optimized, resulting in a more discerning periodic action representation. Our approach significantly reduces the impact of noise, thereby improving counting accuracy. Additionally, the framework is designed to be scalable and adaptable to different types of video content. Experimental results on the RepCountA and UCFRep datasets demonstrate that our proposed method effectively handles repetitive action counting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。