分析深度音乐转录模型在不同音乐类型下的性能退化,揭示数据偏见问题。
Sound and Music Biases in Deep Music Transcription Models: A Systematic Analysis
- 构建MDS数据集,测试风格、动态、多声部等音乐变化对模型影响
- 发现音乐类型导致14%的注记准确率下降,音量变化导致20%下降
- 强调和谐结构评估能更好暴露模型弱点,适合研究泛化性的人看
自动音乐转录(AMT)——将音乐音频转换为音符表示——在深度学习推动下快速发展。由于高质量标注音乐数据稀缺,现有进展主要集中在古典钢琴音乐及少数特定数据集。这些系统能否有效泛化到其他音乐情境仍不清楚。本文针对音乐维度的分布偏移(如体裁、力度、多声部程度),引入MDS语料库,包含三个子集:(1) 体裁,(2) 随机,(3) MAEtest,用于模拟不同分布偏移。我们使用传统信息检索与音乐相关评价指标,评估多个先进AMT系统在该语料库上的表现。结果表明,在特定分布偏移下性能显著下降:音符级F1分数因声音条件下降20个百分点,因体裁变化下降14个百分点。总体而言,力度估计比起始时刻预测更易受音乐变化影响。基于音乐结构的评价指标有助于识别潜在因素。此外,对随机生成非音乐序列的实验显示,系统在极端分布偏移下存在明显局限。这些发现为深度AMT系统中持续存在的语料库偏差问题提供了新证据。
原文摘要 · Abstract (English)
Automatic Music Transcription (AMT) -- the task of converting music audio into note representations -- has seen rapid progress, driven largely by deep learning systems. Due to the limited availability of richly annotated music datasets, much of the progress in AMT has been concentrated on classical piano music, and even a few very specific datasets. Whether these systems can generalize effectively to other musical contexts remains an open question. Complementing recent studies on distribution shifts in sound (e.g., recording conditions), in this work we investigate the musical dimension -- specifically, variations in genre, dynamics, and polyphony levels. To this end, we introduce the MDS corpus, comprising three distinct subsets -- (1) Genre, (2) Random, and (3) MAEtest -- to emulate different axes of distribution shift. We evaluate the performance of several state-of-the-art AMT systems on the MDS corpus using both traditional information-retrieval and musically-informed performance metrics. Our extensive evaluation isolates and exposes varying degrees of performance degradation under specific distribution shifts. In particular, we measure a note-level F1 performance drop of 20 percentage points due to sound, and 14 due to genre. Generally, we find that dynamics estimation proves more vulnerable to musical variation than onset prediction. Musically informed evaluation metrics, particularly those capturing harmonic structure, help identify potential contributing factors. Furthermore, experiments with randomly generated, non-musical sequences reveal clear limitations in system performance under extreme musical distribution shifts. Altogether, these findings offer new evidence of the persistent impact of the Corpus Bias problem in deep AMT systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。