用音频定位引导的混合方法,提升少标签视频动作识别性能
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
- 设计音频源定位引导的跨模态混合策略,利用音视频关系增强学习
- 在UCF-51、Kinetics-400和VGGSound上实现优于现有方法的准确率
- 适合多模态视频理解与低资源标注场景的研究者参考
视频动作识别是理解视频内容的关键任务,但标注成本高。半监督学习(SSL)可在少量标注数据下提升性能,但以往研究多仅依赖视觉模态。视频具有多模态特性,结合音视频信息可进一步提升表现,但该方向尚未充分探索。为此,本文提出一种音视频联合的半监督动作识别框架,即使在极少标注数据下仍具挑战性。为最大化音视频信息,提出一种新型音频源定位引导的Mixup方法,显式建模跨模态关系。在UCF-51、Kinetics-400和VGGSound数据集上的实验表明,所提方法显著优于现有基线,验证了框架与Mixup策略的有效性。
原文摘要 · Abstract (English)
Video action recognition is a challenging but important task for understanding and discovering what the video does. However, acquiring annotations for a video is costly, and semi-supervised learning (SSL) has been studied to improve performance even with a small number of labeled data in the task. Prior studies for semi-supervised video action recognition have mostly focused on using single modality - visuals - but the video is multi-modal, so utilizing both visuals and audio would be desirable and improve performance further, which has not been explored well. Therefore, we propose audio-visual SSL for video action recognition, which uses both visual and audio together, even with quite a few labeled data, which is challenging. In addition, to maximize the information of audio and video, we propose a novel audio source localization-guided mixup method that considers inter-modal relations between video and audio modalities. In experiments on UCF-51, Kinetics-400, and VGGSound datasets, our model shows the superior performance of the proposed semi-supervised audio-visual action recognition framework and audio source localization-guided mixup.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。