arXiv:2502.07328cs.SDcs.AI2025-02NAACL被引 16

研究音乐生成模型的跨文化偏见,提出用小数据微调改善非西方音乐表现。

Music for All: Representational Bias and Cross-Cultural Adaptability of Music Generation Models

  • 用数据量化音乐生成模型对非西方音乐的覆盖不足
  • 仅5.7%音乐数据来自非西方传统,导致模型性能差异显著
  • 小规模微调可提升跨文化适应性,但仍有挑战

音乐-语言模型极大提升了AI自动作曲能力,但其覆盖的音乐类型与文化仍存在局限。本文研究了音乐生成领域的数据集与论文,量化了各类别音乐的代表性偏差。结果显示,现有音乐数据集中仅有5.7%来自非西方音乐类型,这直接导致模型在不同音乐风格间表现不均。我们进一步评估了参数高效微调(PEFT)技术在缓解此类偏差中的效果。以MusicGen和Mustango两个主流模型,在印度古典音乐与土耳其马卡姆音乐两类代表性不足的传统音乐上进行实验,表明通过小规模数据微调可实现跨风格迁移,但也揭示了实际应用中的复杂性。研究呼吁构建更具公平性的基准音乐-语言模型,以支持跨文化迁移学习。

原文摘要 · Abstract (English)

The advent of Music-Language Models has greatly enhanced the automatic music generation capability of AI systems, but they are also limited in their coverage of the musical genres and cultures of the world. We present a study of the datasets and research papers for music generation and quantify the bias and under-representation of genres. We find that only 5.7% of the total hours of existing music datasets come from non-Western genres, which naturally leads to disparate performance of the models across genres. We then investigate the efficacy of Parameter-Efficient Fine-Tuning (PEFT) techniques in mitigating this bias. Our experiments with two popular models -- MusicGen and Mustango, for two underrepresented non-Western music traditions -- Hindustani Classical and Turkish Makam music, highlight the promises as well as the non-triviality of cross-genre adaptation of music through small datasets, implying the need for more equitable baseline music-language models that are designed for cross-cultural transfer learning.

音乐生成跨文化偏见缓解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。