arXiv:2607.06986cs.SD2026-07中稿 · Interspeech 2026

构建多音乐风格歌声合成基准,揭示现有模型跨风格泛化能力弱。

MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres

论文配图:MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
图 1 · 摘自论文原文
  • 自动构建风格对齐乐谱,支持10大类26小类风格评测
  • 主流模型合成歌声在不同风格间差异小,难以区分
  • 轻量微调可显著提升跨风格表现,适合音乐生成研究者

歌声合成(SVS)发展迅速,但其在多样音乐风格间的泛化能力仍缺乏系统研究。现有基准严重偏向流行音乐,难以全面分析风格依赖行为。我们提出MMGenre,一个支持自动构建风格对齐乐谱的多风格歌声合成评测基准,覆盖10大类26小类音乐风格,实现风格感知合成的全面评估。对代表性SVS模型的广泛测试表明,不同风格合成歌声的声学特征高度相似,可分性弱。零样本风格迁移仅带来微弱改善,而轻量级风格特定微调则带来显著提升。MMGenre为多风格歌声合成提供标准化评估框架,揭示了实现真正风格感知合成的关键挑战。

原文摘要 · Abstract (English)

Singing voice synthesis (SVS) has progressed rapidly, yet its ability to generalize across diverse musical genres remains underexplored. Existing benchmarks are heavily biased toward pop music, limiting systematic analysis of genre-dependent behavior. We introduce MMGenre, a benchmark for multi-genre SVS diagnosis, supported by an automatic pipeline for constructing genre-aligned music scores. MMGenre spans 10 major genres and 26 subgenres, enabling comprehensive analysis of genre-aware synthesis. Extensive evaluation of representative SVS models reveals limited genre discrimination: synthesized vocals across genres exhibit highly similar acoustic characteristics and weak separability. While zero-shot genre adaptation yields only marginal improvements, lightweight genre-specific continued training leads to substantial gains. MMGenre provides a standardized framework for multi-genre SVS evaluation and exposes critical challenges in achieving genre-aware singing voice synthesis.

歌声合成多风格基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。