arXiv:2507.06670cs.SDeess.AS2025-07ACL被引 9

首个统一解决歌唱标注三大难题的框架,提升音高、对齐与风格标注质量。

STARS: A Unified Framework for Singing Transcription, Alignment, and Refined Style Annotation

  • 分层处理声学特征,非自回归编码实现多粒度表示学习
  • 在多个指标上超越现有方法,音高定位与对齐精度显著提升
  • 适合需要高质量标注数据的歌声合成研究者使用

近期歌声合成(SVS)的突破推动了高质量标注数据集的需求,但人工标注成本高昂。现有自动歌唱标注(ASA)方法多仅解决流程中单一环节。为此,我们提出STARS,据我们所知首个统一处理歌唱转录、对齐与精细风格标注的框架。该框架提供多层次标注:(1)精确的音素-音频对齐,(2)鲁棒的音符转录与时间定位,(3)表现性发声技术识别,(4)包含情绪与节奏的全局风格刻画。其架构采用跨帧、词、音素、音符、句子层级的分层声学特征处理。新颖的非自回归局部声学编码器支持结构化层次表示学习。实验验证表明,该框架在多个评估维度上优于现有方法。此外,在SVS训练中的应用显示,使用STARS标注数据的模型在感知自然度和风格控制精度上均有显著提升。本工作不仅缓解了歌唱数据集构建的可扩展性瓶颈,更开创了可控歌声合成的新范式。音频样例见 https://gwx314.github.io/stars-demo/。

原文摘要 · Abstract (English)

Recent breakthroughs in singing voice synthesis (SVS) have heightened the demand for high-quality annotated datasets, yet manual annotation remains prohibitively labor-intensive and resource-intensive. Existing automatic singing annotation (ASA) methods, however, primarily tackle isolated aspects of the annotation pipeline. To address this fundamental challenge, we present STARS, which is, to our knowledge, the first unified framework that simultaneously addresses singing transcription, alignment, and refined style annotation. Our framework delivers comprehensive multi-level annotations encompassing: (1) precise phoneme-audio alignment, (2) robust note transcription and temporal localization, (3) expressive vocal technique identification, and (4) global stylistic characterization including emotion and pace. The proposed architecture employs hierarchical acoustic feature processing across frame, word, phoneme, note, and sentence levels. The novel non-autoregressive local acoustic encoders enable structured hierarchical representation learning. Experimental validation confirms the framework's superior performance across multiple evaluation dimensions compared to existing annotation approaches. Furthermore, applications in SVS training demonstrate that models utilizing STARS-annotated data achieve significantly enhanced perceptual naturalness and precise style control. This work not only overcomes critical scalability challenges in the creation of singing datasets but also pioneers new methodologies for controllable singing voice synthesis. Audio samples are available at https://gwx314.github.io/stars-demo/.

歌声合成语音标注风格控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。