arXiv:2511.15270cs.SD2025-11

构建了超百万级歌曲改编数据集,支持参考式音乐生成研究。

LargeSHS: A large-scale dataset of music adaptation

  • 基于SecondHandSongs构建,包含170万条元数据与90万音频链接
  • 首次提供结构化歌曲改编关系,可构建覆盖家族与表演聚类
  • 适合研究翻唱生成、参考式音乐创作及适应性音乐信息检索

近年来,基于AI的音乐生成主要聚焦于文本条件模型,而对参考式生成(如歌曲改编)关注较少。为推动该方向研究,我们引入LargeSHS,一个源自SecondHandSongs的大规模数据集,包含超过170万条元数据和约90万条公开可访问的音频链接。与现有数据集不同,LargeSHS包含音乐作品间的结构化改编关系,支持构建改编树与表演聚类,展现翻唱家族结构。我们提供了详尽的统计数据与现有数据集对比,凸显LargeSHS在规模与丰富性上的独特优势。该数据集为翻唱生成、参考式音乐生成及适应性音乐信息检索(MIR)任务开辟新路径。

原文摘要 · Abstract (English)

Recent advances in AI-based music generation have focused heavily on text-conditioned models, with less attention given to reference-based generation such as song adaptation. To support this line of research, we introduce LargeSHS, a large-scale dataset derived from SecondHandSongs, containing over 1.7 million metadata entries and approximately 900k publicly accessible audio links. Unlike existing datasets, LargeSHS includes structured adaptation relationships between musical works, enabling the construction of adaptation trees and performance clusters that represent cover song families. We provide comprehensive statistics and comparisons with existing datasets, highlighting the unique scale and richness of LargeSHS. This dataset paves the way for new research in cover song generation, reference-based music generation, and adaptation-aware MIR tasks.

音乐生成数据集参考式生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。