首个跨语言多模态歌词翻译数据集,助力动画歌曲精准翻唱
MAVL: A Multilingual Audio-Video Lyrics Dataset for Animated Song Translation
- 结合音视频线索与音节数约束,生成可演唱的译文
- 在可唱性与语义准确率上显著优于纯文本模型
- 适合音乐翻译、动画字幕、多模态生成研究者
歌词翻译需兼顾语义准确传递与音乐节奏、音节结构和诗学风格的保留。在动画音乐剧场景中,还需对齐视听线索,挑战更大。本文提出首个跨语言、多模态的动画歌曲翻译基准数据集MAVL,融合文本、音频与视频信息,支持更丰富、更具表现力的译文生成。基于此,我们构建了音节数约束的音视频大模型SylAVL-CoT,通过链式思维推理,利用音视频线索并强制音节匹配,生成自然流畅的可演唱歌词。实验表明,SylAVL-CoT在可唱性与上下文准确性上显著优于纯文本模型,凸显多模态、跨语言方法在歌词翻译中的价值。
原文摘要 · Abstract (English)
Lyrics translation requires both accurate semantic transfer and preservation of musical rhythm, syllabic structure, and poetic style. In animated musicals, the challenge intensifies due to alignment with visual and auditory cues. We introduce Multilingual Audio-Video Lyrics Benchmark for Animated Song Translation (MAVL), the first multilingual, multimodal benchmark for singable lyrics translation. By integrating text, audio, and video, MAVL enables richer and more expressive translations than text-only approaches. Building on this, we propose Syllable-Constrained Audio-Video LLM with Chain-of-Thought SylAVL-CoT, which leverages audio-video cues and enforces syllabic constraints to produce natural-sounding lyrics. Experimental results demonstrate that SylAVL-CoT significantly outperforms text-based models in singability and contextual accuracy, emphasizing the value of multimodal, multilingual approaches for lyrics translation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。