让音乐模型理解全球多元文化,提升非西方音乐识别效果。
CultureMERT: Continual Pre-Training for Cross-Cultural Music Representation Learning
- 采用两阶段持续预训练,动态调整学习率以稳定跨文化适应。
- 在650小时多文化数据上训练,非西方音乐标签任务平均提升4.9%。
- 单文化模型融合法(任务算术)效果接近多文化模型,且不丢西式数据表现。
近期音乐基础模型虽提升了音频表征学习能力,但在多样音乐传统中的有效性仍受限。本文提出CultureMERT-95M,一种多文化适配的基础模型,旨在增强跨文化音乐表征学习与理解。为此,我们设计了一种两阶段持续预训练策略,结合学习率重暖与重衰减机制,实现资源受限下的稳定适应。在包含希腊、土耳其和印度音乐传统的650小时多文化数据集上训练后,该模型在多种非西方音乐自动标注任务中平均提升ROC-AUC和AP达4.9%,超越现有最优方法,且对西式基准无明显遗忘。进一步研究显示,任务算术(task arithmetic)——通过权重空间合并单文化适配模型——在非西方任务上表现相当,且在西式数据上无性能下降。跨文化评估表明,单文化模型跨文化迁移效果不一,而多文化适配模型整体表现最佳。为推动世界音乐表征研究,我们公开发布CultureMERT-95M与CultureMERT-TA-95M,助力更具有文化敏感性的音乐基础模型发展。
原文摘要 · Abstract (English)
Recent advances in music foundation models have improved audio representation learning, yet their effectiveness across diverse musical traditions remains limited. We introduce CultureMERT-95M, a multi-culturally adapted foundation model developed to enhance cross-cultural music representation learning and understanding. To achieve this, we propose a two-stage continual pre-training strategy that integrates learning rate re-warming and re-decaying, enabling stable adaptation even with limited computational resources. Training on a 650-hour multi-cultural data mix, comprising Greek, Turkish, and Indian music traditions, results in an average improvement of 4.9% in ROC-AUC and AP across diverse non-Western music auto-tagging tasks, surpassing prior state-of-the-art, with minimal forgetting on Western-centric benchmarks. We further investigate task arithmetic, an alternative approach to multi-cultural adaptation that merges single-culture adapted models in the weight space. Task arithmetic performs on par with our multi-culturally trained model on non-Western auto-tagging tasks and shows no regression on Western datasets. Cross-cultural evaluation reveals that single-culture models transfer with varying effectiveness across musical traditions, whereas the multi-culturally adapted model achieves the best overall performance. To support research on world music representation learning, we publicly release CultureMERT-95M and CultureMERT-TA-95M, fostering the development of more culturally aware music foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。