自动化处理歌曲数据,一键生成可用于训练的结构化歌词与时间戳。
SongPrep: A Preprocessing Framework and End-to-end Model for Full-song Structure Parsing and Lyrics Transcription
- 构建端到端框架,自动完成分轨、结构分析与歌词识别。
- 在SSLD-200数据集上实现低误辨率(DER)与低词错误率(WER)。
- 适合想高效训练高质量音乐生成模型的研究者使用。
人工智能生成内容(AIGC)是当前热门研究方向,其中歌曲生成备受关注。尽管歌曲资源丰富,但有效数据准备仍是难题,传统方法需大量人工标注,耗时且成本高。为此,我们提出SongPrep,一个专为歌曲数据设计的自动化预处理流程,可自动完成源分离、结构分析和歌词识别,生成可直接用于训练的结构化数据。此外,我们还推出基于预训练语言模型的端到端结构化歌词识别模型SongPrepE2E,无需额外源分离即可分析整首歌的结构与歌词,并提供精确时间戳。利用全曲上下文与预训练语义知识,SongPrepE2E在新提出的SSLD-200数据集上取得低误辨率(DER)与低词错误率(WER)。下游任务表明,使用SongPrepE2E输出的数据训练的歌曲生成模型,其生成结果与人类创作高度相似。
原文摘要 · Abstract (English)
Artificial Intelligence Generated Content (AIGC) is currently a popular research area. Among its various branches, song generation has attracted growing interest. Despite the abundance of available songs, effective data preparation remains a significant challenge. Converting these songs into training-ready datasets typically requires extensive manual labeling, which is both time consuming and costly. To address this issue, we propose SongPrep, an automated preprocessing pipeline designed specifically for song data. This framework streamlines key processes such as source separation, structure analysis, and lyric recognition, producing structured data that can be directly used to train song generation models. Furthermore, we introduce SongPrepE2E, an end-to-end structured lyrics recognition model based on pretrained language models. Without the need for additional source separation, SongPrepE2E is able to analyze the structure and lyrics of entire songs and provide precise timestamps. By leveraging context from the whole song alongside pretrained semantic knowledge, SongPrepE2E achieves low Diarization Error Rate (DER) and Word Error Rate (WER) on the proposed SSLD-200 dataset. Downstream tasks demonstrate that training song generation models with the data output by SongPrepE2E enables the generated songs to closely resemble those produced by humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。