用双重引导提升动作检测伪框质量,标签少时效果更优
Dual Guidance Semi-Supervised Action Detection
- 结合帧级分类与框预测,双重约束伪框生成
- 在少量标注数据下,多个数据集上显著提升性能
- 适合弱监督动作定位场景,尤其标签稀缺时
半监督学习(SSL)在标注难获取的场景中展现出巨大潜力。然而,现有研究主要集中在图像分类任务。本文提出一种时空动作定位的半监督方法,设计双引导网络以筛选更优的伪边界框。该方法融合帧级分类与边界框预测,强制动作类别在帧间和框内保持一致性。在UCF101-24、J-HMDB-21和AVA等多个知名时空动作定位数据集上的评估表明,所提模块在标注数据有限的情况下显著提升了模型表现,优于扩展的图像类半监督基线方法。
原文摘要 · Abstract (English)
Semi-Supervised Learning (SSL) has shown tremendous potential to improve the predictive performance of deep learning models when annotations are hard to obtain. However, the application of SSL has so far been mainly studied in the context of image classification. In this work, we present a semi-supervised approach for spatial-temporal action localization. We introduce a dual guidance network to select better pseudo-bounding boxes. It combines a frame-level classification with a bounding-box prediction to enforce action class consistency across frames and boxes. Our evaluation across well-known spatial-temporal action localization datasets, namely UCF101-24 , J-HMDB-21 and AVA shows that the proposed module considerably enhances the model's performance in limited labeled data settings. Our framework achieves superior results compared to extended image-based semi-supervised baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。