arXiv:2602.24071cs.SDcs.CL2026-02AAAI被引 4

首个能生成宋词古曲的模型,还原千年诗词韵律之美。

SongSong: A Time Phonograph for Chinese SongCi Music from Thousand of Years Away

  • 分步生成:先定旋律,再合成唱腔与伴奏
  • 自建29.9小时宋词音乐数据集,填补古乐空白
  • 评测胜过Suno等平台,适合古风音乐创作

近年来,音乐生成技术发展迅速,但现有模型多聚焦现代流行歌曲,难以复现如宋词般具有独特节奏与风格的古代音乐。本文提出SongSong,首个可恢复中国宋词音乐的生成模型。该模型首先根据输入宋词预测旋律,再分别生成演唱声线与伴奏,并融合成完整乐曲。为解决古代音乐数据稀缺问题,我们构建了包含29.9小时作品的OpenSongSong数据集,涵盖多位宋代词乐大师之作。通过在85句未参与训练的宋词上进行主观与客观评估,结果表明SongSong在生成质量上优于Suno、SkyMusic等平台,展现出领先性能。

原文摘要 · Abstract (English)

Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient Chinese SongCi. In this paper, we introduce SongSong, the first music generation model capable of restoring Chinese SongCi to our knowledge. Our model first predicts the melody from the input SongCi, then separately generates the singing voice and accompaniment based on that melody, and finally combines all elements to create the final piece of music. Additionally, to address the lack of ancient music datasets, we create OpenSongSong, a comprehensive dataset of ancient Chinese SongCi music, featuring 29.9 hours of compositions by various renowned SongCi music masters. To assess SongSong's proficiency in performing SongCi, we randomly select 85 SongCi sentences that were not part of the training set for evaluation against SongSong and music generation platforms such as Suno and SkyMusic. The subjective and objective outcomes indicate that our proposed model achieves leading performance in generating high-quality SongCi music.

音乐生成宋词古风音频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。