arXiv:2605.04613cs.SDcs.AI2026-05

用大模型统一生成歌声歌词旋律,自动标注更准更快。

VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models

论文配图:VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
图 1 · 摘自论文原文
  • 用交错提示联合建模歌词、旋律和音符对应关系。
  • 在多个数据集上达到当前最优转录效果。
  • 适合需要高质量歌声标注的音乐生成研究者。

高质量的歌声标注是现代歌唱语音合成(SVS)系统的基础。然而,由于人工标注需大量人力和音乐专业知识,大规模标注不现实,因此自动标注至关重要。尽管如此,现有自动转录系统仍面临显著挑战:依赖复杂的多阶段流程,难以恢复文本-音符对齐,且对分布外(OOD)歌声数据泛化能力差。为此,我们提出VocalParse,一个基于大音频语言模型(LALM)的统一歌声转录(SVT)模型。创新性地引入交错提示机制,联合建模歌词、旋律与词-音符对应关系,生成直接映射到结构化乐谱的序列。此外,提出类思维链(CoT)提示策略,先解码歌词作为语义框架,显著缓解上下文干扰问题,同时保留交错生成的结构优势。实验表明,VocalParse在多个歌唱数据集上实现当前最优的转录性能。源代码与检查点已公开于 https://github.com/pymaster17/VocalParse。

原文摘要 · Abstract (English)

High-quality singing annotations are fundamental to modern Singing Voice Synthesis (SVS) systems. However, obtaining these annotations at scale through manual labeling is unrealistic due to the substantial labor and musical expertise required, making automatic annotation highly necessary. Despite their utility, current automatic transcription systems face significant challenges: they often rely on complex multi-stage pipelines, struggle to recover text-note alignments, and exhibit poor generalization to out-of-distribution (OOD) singing data. To alleviate these issues, we present VocalParse, a unified singing voice transcription (SVT) model built upon a Large Audio Language Model (LALM). Specifically, our novel contribution is to introduce an interleaved prompting formulation that jointly models lyrics, melody, and word-note correspondence, yielding a generated sequence that directly maps to a structured musical score. Furthermore, we propose a Chain-of-Thought (CoT) style prompting strategy, which decodes lyrics first as a semantic scaffold, significantly mitigating the context disruption problem while preserving the structural benefits of interleaved generation. Experiments demonstrate that VocalParse achieves state-of-the-art SVT performance on multiple singing datasets. The source code and checkpoint are available at https://github.com/pymaster17/VocalParse.

歌声转录大模型音频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。