用场景信息生成部分标签,降低音效标注成本。
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
- 利用声学场景构建音效的可能标签集合
- 半监督框架融合强标签与部分标签提升性能
- 自蒸馏方法优化部分标签,改善模型训练
标注音效的时间边界耗时费力,制约了强监督学习在音频检测中的可扩展性。为降低标注成本,已有研究采用仅需片段级标签的弱监督学习。作为替代方案,部分标签学习提供了一种低成本方法:给出一组可能的标签而非精确的弱标注。然而,音频分析中的部分标签学习仍鲜有探索。受声学场景能为音效提供上下文信息的启发,本文提出一种多任务学习框架,联合进行声学场景分类与带部分标签的音效检测。该方法在降低标注成本的同时,通过引入半监督机制融合强标签与部分标签,缓解因缺乏精确事件集和时间标注导致的检测性能下降问题。此外,还提出基于自蒸馏的标签精炼方法,以进一步优化部分标签并提升模型训练效果。
原文摘要 · Abstract (English)
Annotating time boundaries of sound events is labor-intensive, limiting the scalability of strongly supervised learning in audio detection. To reduce annotation costs, weakly-supervised learning with only clip-level labels has been widely adopted. As an alternative, partial label learning offers a cost-effective approach, where a set of possible labels is provided instead of exact weak annotations. However, partial label learning for audio analysis remains largely unexplored. Motivated by the observation that acoustic scenes provide contextual information for constructing a set of possible sound events, we utilize acoustic scene information to construct partial labels of sound events. On the basis of this idea, in this paper, we propose a multitask learning framework that jointly performs acoustic scene classification and sound event detection with partial labels of sound events. While reducing annotation costs, weakly-supervised and partial label learning often suffer from decreased detection performance due to lacking the precise event set and their temporal annotations. To better balance between annotation cost and detection performance, we also explore a semi-supervised framework that leverages both strong and partial labels. Moreover, to refine partial labels and achieve better model training, we propose a label refinement method based on self-distillation for the proposed approach with partial labels.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。