自监督多音高估计中发现模型过拟合与性能退化现象。
Investigating an Overfitting and Degeneration Phenomenon in Self-Supervised Multi-Pitch Estimation
- 结合音高不变/等变性设计自监督目标,联合训练提升性能。
- 在更大规模数据上训练时,模型对监督数据过拟合,自监督数据表现退化。
- 揭示了自监督学习中的潜在冲突机制,适合关注模型可靠性研究者阅读。
多音高估计(MPE)是音乐信息检索(MIR)系统的重要能力,对乐谱转录等下游任务至关重要。然而现有方法主要依赖有监督学习,标注数据收集困难。近期基于音高和谐波信号内在特性的自监督技术在单音和多音高估计中展现出潜力,但仍逊于有监督方法。本文扩展经典有监督MPE范式,引入基于音高不变性和音高等变性的多个自监督目标进行联合训练,在封闭训练条件下取得显著提升,暗示在更大规模数据上应用相同目标可进一步优化。然而,实际应用中我们发现模型同时对有监督数据过拟合,且在仅用于自监督的数据上性能退化。本文揭示并分析了这一现象,提出了对根本问题的见解。
原文摘要 · Abstract (English)
Multi-Pitch Estimation (MPE) continues to be a sought after capability of Music Information Retrieval (MIR) systems, and is critical for many applications and downstream tasks involving pitch, including music transcription. However, existing methods are largely based on supervised learning, and there are significant challenges in collecting annotated data for the task. Recently, self-supervised techniques exploiting intrinsic properties of pitch and harmonic signals have shown promise for both monophonic and polyphonic pitch estimation, but these still remain inferior to supervised methods. In this work, we extend the classic supervised MPE paradigm by incorporating several self-supervised objectives based on pitch-invariant and pitch-equivariant properties. This joint training results in a substantial improvement under closed training conditions, which naturally suggests that applying the same objectives to a broader collection of data will yield further improvements. However, in doing so we uncover a phenomenon whereby our model simultaneously overfits to the supervised data while degenerating on data used for self-supervision only. We demonstrate and investigate this and offer our insights on the underlying problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。