arXiv:2409.01448cs.CVcs.LG2024-09ECCV被引 6

用时间对齐度量提升细粒度动作识别的伪标签质量

FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition

  • 引入可学习的对齐可验证度量,捕捉动作阶段特征
  • 在4个细粒度数据集上显著超越现有方法
  • 适合需要精准动作分析的应用场景

真实场景中的动作识别常需理解细微动作差异,如体育分析、AR/VR交互和手术视频。尽管细粒度动作标注成本高,现有半监督方法多聚焦粗粒度动作识别。由于缺乏场景偏差,细粒度动作分类需依赖动作阶段理解,而现有粗粒度方法效果不佳。本文首次系统研究半监督细粒度动作识别(FGAR),发现动态时间规整(DTW)等对齐距离能有效衡量动作阶段,但传统DTW为成对且严格对齐,不适用于分类。为此,提出基于对齐可验证性的度量学习方法,构建可学习的对齐可度量分数,用于优化主视频编码器的伪标签。所提协同伪标签框架FinePseudo在四个细粒度数据集(Diving48、FineGym99、FineGym288、FineDiving)上显著优于基线,并在Kinetics400和Something-SomethingV2等粗粒度数据集上也取得提升。实验还验证了其在开放世界设置下处理新未标注类别的鲁棒性。

原文摘要 · Abstract (English)

Real-life applications of action recognition often require a fine-grained understanding of subtle movements, e.g., in sports analytics, user interactions in AR/VR, and surgical videos. Although fine-grained actions are more costly to annotate, existing semi-supervised action recognition has mainly focused on coarse-grained action recognition. Since fine-grained actions are more challenging due to the absence of scene bias, classifying these actions requires an understanding of action-phases. Hence, existing coarse-grained semi-supervised methods do not work effectively. In this work, we for the first time thoroughly investigate semi-supervised fine-grained action recognition (FGAR). We observe that alignment distances like dynamic time warping (DTW) provide a suitable action-phase-aware measure for comparing fine-grained actions, a concept previously unexploited in FGAR. However, since regular DTW distance is pairwise and assumes strict alignment between pairs, it is not directly suitable for classifying fine-grained actions. To utilize such alignment distances in a limited-label setting, we propose an Alignability-Verification-based Metric learning technique to effectively discriminate between fine-grained action pairs. Our learnable alignability score provides a better phase-aware measure, which we use to refine the pseudo-labels of the primary video encoder. Our collaborative pseudo-labeling-based framework `\textit{FinePseudo}' significantly outperforms prior methods on four fine-grained action recognition datasets: Diving48, FineGym99, FineGym288, and FineDiving, and shows improvement on existing coarse-grained datasets: Kinetics400 and Something-SomethingV2. We also demonstrate the robustness of our collaborative pseudo-labeling in handling novel unlabeled classes in open-world semi-supervised setups. Project Page: https://daveishan.github.io/finepsuedo-webpage/.

细粒度识别伪标签动作分析半监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。