用一个音乐基础模型的中间特征,提升多种下游任务表现。
Music Foundation Model as Generic Booster for Music Downstream Tasks
- 用层次化特征提取器增强音乐任务表征能力
- 在标签、转录等任务上显著提升性能
- 特别适合数据少的任务场景
我们验证了单一音乐基础模型(MFM)提取的中间表示在多个音乐下游任务中的有效性。提出SoniDo,一种可从目标音乐样本中提取分层特征的音乐基础模型。通过利用分层中间特征,SoniDo控制信息粒度,从而在音乐标签、音乐转录、音源分离和音乐混音等典型任务上均取得性能提升。实验表明,基础模型提取的特征能有效增强下游模型训练效果,证明其作为通用助推器的潜力。该方法不仅优化现有任务模型,还适用于数据稀缺的音乐下游任务,为更高效、易用的音乐处理方案提供新路径。
原文摘要 · Abstract (English)
We demonstrate the efficacy of using intermediate representations from a single foundation model to enhance various music downstream tasks. We introduce SoniDo, a music foundation model (MFM) designed to extract hierarchical features from target music samples. By leveraging hierarchical intermediate features, SoniDo constrains the information granularity, leading to improved performance across various downstream tasks including both understanding and generative tasks. We specifically evaluated this approach on representative tasks such as music tagging, music transcription, music source separation, and music mixing. Our results reveal that the features extracted from foundation models provide valuable enhancements in training downstream task models. This highlights the capability of using features extracted from music foundation models as a booster for downstream tasks. Our approach not only benefits existing task-specific models but also supports music downstream tasks constrained by data scarcity. This paves the way for more effective and accessible music processing solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。