首个波斯音乐生成数据集,让模型学会创作正宗波斯乐曲。
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
- 构建900小时波斯音乐数据集,覆盖流行、传统等多风格。
- 微调MusicGen后,生成音乐更贴合波斯风格标签。
- 为非西方音乐生成提供可复用的文化适配范例。
波斯音乐以其独特的音高体系、调式系统(Dastgah)和节奏结构,对主要基于西方音乐训练的生成模型构成挑战。为此,我们首次构建了大规模波斯歌曲数据集,包含超过900小时高质量音频样本,涵盖流行、传统和当代等多种子风格,全面捕捉波斯音乐的旋律与文化多样性。该数据集成为微调前沿生成模型MusicGen的基础。通过主观与客观指标评估模型性能,我们报告了生成音乐在语义上与预期风格标签的一致性比例。结果表明,微调后的模型生成作品更符合波斯音乐风格规范。本工作不仅引入了一个新的生成音乐研究资源,也展示了音乐生成模型在非主流文化语境下的适应能力。
原文摘要 · Abstract (English)
Persian music, with its unique tonalities, modal systems (Dastgah), and rhythmic structures, presents significant challenges for music generation models trained primarily on Western music. We address this gap by curating the first large-scale dataset of Persian songs, comprising over 900 hours high-quality audio samples across diverse sub-genres, including pop, traditional, and contemporary styles. This dataset captures the rich melodic and cultural diversity of Persian music and serves as the foundation for fine-tuning MusicGen, a state-of-the-art generative music model. We adapt MusicGen to this domain and evaluate its performance by utilizing subjective and objective metrics. To assess the semantic alignment between generated music and intended style tags, we report the proportion of relevant tags accurately reflected in the generated outputs. Our results demonstrate that the fine-tuned model produces compositions that more align with Persian stylistic conventions. This work introduces a new resource for generative music research and illustrates the adaptability of music generation models to underrepresented cultural and linguistic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。