用时间扰动与学习稳定化提升细粒度动作识别的半监督性能
SeFAR: Semi-supervised Fine-grained Action Recognition with Temporal Perturbation and Learning Stabilization
- 构建双层次时间特征,结合适度时间扰动增强教师-学生模型
- 在FineGym和FineDiving上达当前最佳,跨数据规模表现稳健
- 适合需要细粒度动作理解的多模态系统研发者参考
人体动作理解对多模态系统发展至关重要。尽管近期基于大语言模型(LLMs)的方法趋向通用化,却常忽视特定能力需求。本文聚焦更具挑战性的细粒度动作识别(FAR),其关注短时程内精细语义标签(如“后空翻屈体转一圈”)。鉴于标注细粒度标签成本高且微调LLMs需大量数据,我们采用半监督学习(SSL)。提出SeFAR框架,通过构建双层次时间元素,设计含适度时间扰动的强增强策略,用于教师-学生学习范式。为应对教师模型预测中高不确定性,引入自适应调节机制以稳定学习过程。实验表明,SeFAR在两个FAR数据集FineGym与FineDiving上均达当前最优,在不同数据规模下表现稳定;同时在经典粗粒度数据集UCF101与HMDB51上优于其他半监督方法。进一步分析与消融实验验证了各设计的有效性。此外,所提取特征显著提升多模态基础模型对细粒度、领域特定语义的理解能力。
原文摘要 · Abstract (English)
Human action understanding is crucial for the advancement of multimodal systems. While recent developments, driven by powerful large language models (LLMs), aim to be general enough to cover a wide range of categories, they often overlook the need for more specific capabilities. In this work, we address the more challenging task of Fine-grained Action Recognition (FAR), which focuses on detailed semantic labels within shorter temporal duration (e.g., "salto backward tucked with 1 turn"). Given the high costs of annotating fine-grained labels and the substantial data needed for fine-tuning LLMs, we propose to adopt semi-supervised learning (SSL). Our framework, SeFAR, incorporates several innovative designs to tackle these challenges. Specifically, to capture sufficient visual details, we construct Dual-level temporal elements as more effective representations, based on which we design a new strong augmentation strategy for the Teacher-Student learning paradigm through involving moderate temporal perturbation. Furthermore, to handle the high uncertainty within the teacher model's predictions for FAR, we propose the Adaptive Regulation to stabilize the learning process. Experiments show that SeFAR achieves state-of-the-art performance on two FAR datasets, FineGym and FineDiving, across various data scopes. It also outperforms other semi-supervised methods on two classical coarse-grained datasets, UCF101 and HMDB51. Further analysis and ablation studies validate the effectiveness of our designs. Additionally, we show that the features extracted by our SeFAR could largely promote the ability of multimodal foundation models to understand fine-grained and domain-specific semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。