用指令控制语音节奏与情感,实现零样本电影配音的精准对口型。
InstructDubber: Instruction-based Alignment for Zero-shot Movie Dubbing
- 通过大语言模型生成自然语言指令,指导语音节奏与情感。
- 在三个基准上均超越现有方法,零样本场景下提升显著。
- 无需复杂视觉处理,适合跨域配音,适合影视字幕自动化场景。
电影配音旨在根据剧本合成特定声线的语音,同时保证口型同步与情感语调与角色视觉表演一致。然而,现有基于视觉特征的对齐方法存在两大局限:(1)依赖复杂的手工视觉预处理流程,包括面部关键点检测与特征提取;(2)对未见视觉领域泛化能力差,常导致对齐质量下降。为此,我们提出 InstructDubber,一种基于指令的对齐配音方法,适用于鲁棒的域内与零样本电影配音。具体而言,首先将视频、剧本及对应提示输入多模态大语言模型,生成关于说话速率与情绪状态的自然语言配音指令,该方法对视觉域变化具有鲁棒性。其次,设计指令式时长蒸馏模块,从说话速率指令中挖掘判别性时长线索,预测唇动对齐的音素级发音时长。第三,针对情感-语调对齐,提出指令式情绪校准模块,使用真实配音情绪监督微调基于大语言模型的指令分析器,并据此预测语调。最后,将预测的时长与语调结合剧本输入音频解码器,生成视频对齐的配音。在三个主流基准上的大量实验表明,InstructDubber 在域内与零样本场景下均优于现有最先进方法。
原文摘要 · Abstract (English)
Movie dubbing seeks to synthesize speech from a given script using a specific voice, while ensuring accurate lip synchronization and emotion-prosody alignment with the character's visual performance. However, existing alignment approaches based on visual features face two key limitations: (1)they rely on complex, handcrafted visual preprocessing pipelines, including facial landmark detection and feature extraction; and (2) they generalize poorly to unseen visual domains, often resulting in degraded alignment and dubbing quality. To address these issues, we propose InstructDubber, a novel instruction-based alignment dubbing method for both robust in-domain and zero-shot movie dubbing. Specifically, we first feed the video, script, and corresponding prompts into a multimodal large language model to generate natural language dubbing instructions regarding the speaking rate and emotion state depicted in the video, which is robust to visual domain variations. Second, we design an instructed duration distilling module to mine discriminative duration cues from speaking rate instructions to predict lip-aligned phoneme-level pronunciation duration. Third, for emotion-prosody alignment, we devise an instructed emotion calibrating module, which finetunes an LLM-based instruction analyzer using ground truth dubbing emotion as supervision and predicts prosody based on the calibrated emotion analysis. Finally, the predicted duration and prosody, together with the script, are fed into the audio decoder to generate video-aligned dubbing. Extensive experiments on three major benchmarks demonstrate that InstructDubber outperforms state-of-the-art approaches across both in-domain and zero-shot scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。