arXiv:2510.22295cs.AIcs.CL2025-10

首个越南语歌词自动转录数据集,解决方言与声调难题。

VietLyrics: A Large-Scale Dataset and Models for Vietnamese Automatic Lyrics Transcription

  • 构建647小时越南语音乐数据集,含逐行对齐歌词
  • 微调Whisper模型后性能超越现有多语言系统
  • 适合低资源语言音乐研究者使用

越南语音乐的自动歌词转录(ALT)因声调复杂和方言差异面临独特挑战,但因缺乏专用数据集而研究有限。为此,我们构建了首个大规模越南语ALT数据集VietLyrics,包含647小时歌曲,具备逐行对齐的歌词与元数据。评估现有基于ASR的方法发现其存在频繁错误和非人声段落的幻觉问题。通过在VietLyrics上微调Whisper模型,性能显著优于现有多语言系统(如LyricWhiz)。我们公开发布VietLyrics及模型,旨在推动越南语音乐计算研究,并展示该方法在低资源语言音乐中的潜力。

原文摘要 · Abstract (English)

Automatic Lyrics Transcription (ALT) for Vietnamese music presents unique challenges due to its tonal complexity and dialectal variations, but remains largely unexplored due to the lack of a dedicated dataset. Therefore, we curated the first large-scale Vietnamese ALT dataset (VietLyrics), comprising 647 hours of songs with line-level aligned lyrics and metadata to address these issues. Our evaluation of current ASRbased approaches reveal significant limitations, including frequent transcription errors and hallucinations in non-vocal segments. To improve performance, we fine-tuned Whisper models on the VietLyrics dataset, achieving superior results compared to existing multilingual ALT systems, including LyricWhiz. We publicly release VietLyrics and our models, aiming to advance Vietnamese music computing research while demonstrating the potential of this approach for ALT in low-resource language and music.

歌词转录越南语语音识别数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。