用一致性损失提升歌声混音中的歌词识别准确率
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
- 通过一致性损失对齐人声与混音编码器表征
- 在混音环境下相比基线提升3.2%的字符准确率
- 适合音乐信息处理、语音识别方向的研究者
自动歌词转录(ALT)旨在从演唱声音中识别歌词,类似于语音识别(ASR),但因歌唱语音的领域特性而更具挑战性。尽管基础ASR模型在多种语音任务中表现稳健,但在伴有伴奏的歌唱语音上性能显著下降。本文聚焦这一性能差距,探索使用低秩适配(LoRA)进行ALT,对比单领域与双领域微调策略。提出采用一致性损失,更好对齐人声与混音编码器表示,从而在不依赖人声分离的前提下提升混合音频上的转录效果。实验表明,朴素的双领域微调表现不佳,而加入一致性损失的结构化训练带来稳定但适度的性能提升,在混合场景下字符准确率提升3.2%,验证了将ASR基础模型适配于音乐任务的潜力。
原文摘要 · Abstract (English)
Automatic Lyrics Transcription (ALT) aims to recognize lyrics from singing voices, similar to Automatic Speech Recognition (ASR) for spoken language, but faces added complexity due to domain-specific properties of the singing voice. While foundation ASR models show robustness in various speech tasks, their performance degrades on singing voice, especially in the presence of musical accompaniment. This work focuses on this performance gap and explores Low-Rank Adaptation (LoRA) for ALT, investigating both single-domain and dual-domain fine-tuning strategies. We propose using a consistency loss to better align vocal and mixture encoder representations, improving transcription on mixture without relying on singing voice separation. Our results show that while naïve dual-domain fine-tuning underperforms, structured training with consistency loss yields modest but consistent gains, demonstrating the potential of adapting ASR foundation models for music.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。