arXiv:2412.08988cs.SDcs.MM2024-12CVPR被引 31

让配音既对口型又可控情绪,还能保持发音清晰。

EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

  • 用嘴型与语调的对比学习实现精准对口型
  • 融合视频级音素提升发音可懂度,同步率超95%
  • 通过正负引导机制实现情绪强度和类型精准控制

给定一段文字、视频片段和参考音频,电影配音任务旨在生成与视频对齐且克隆目标语音的语音。现有方法存在两大缺陷:(1)难以同时保证音视频同步与清晰发音;(2)无法表达用户定义的情绪。为此,我们提出 EmoDubber,一种可调控情绪的配音架构,支持用户指定情绪类型和强度,同时实现高质量口型同步与发音清晰。首先设计唇部相关韵律对齐(LPA),通过时长层级的对比学习建模唇动与韵律变化的一致性,增强对齐合理性。其次采用发音增强(PE)策略,利用高效Conformer融合视频级音素序列,提升语音可懂度。再通过说话人身份适配模块解码声学先验并注入说话人风格嵌入。最后,提出的基于流的用户情绪控制(FUEC)通过流匹配预测网络生成波形,其梯度方向与引导尺度由用户情绪指令决定,借助正负引导机制放大目标情绪、抑制其他情绪。在三个基准数据集上的大量实验表明,该方法优于多个当前最优方法。

原文摘要 · Abstract (English)

Given a piece of text, a video clip, and a reference audio, the movie dubbing task aims to generate speech that aligns with the video while cloning the desired voice. The existing methods have two primary deficiencies: (1) They struggle to simultaneously hold audio-visual sync and achieve clear pronunciation; (2) They lack the capacity to express user-defined emotions. To address these problems, we propose EmoDubber, an emotion-controllable dubbing architecture that allows users to specify emotion type and emotional intensity while satisfying high-quality lip sync and pronunciation. Specifically, we first design Lip-related Prosody Aligning (LPA), which focuses on learning the inherent consistency between lip motion and prosody variation by duration level contrastive learning to incorporate reasonable alignment. Then, we design Pronunciation Enhancing (PE) strategy to fuse the video-level phoneme sequences by efficient conformer to improve speech intelligibility. Next, the speaker identity adapting module aims to decode acoustics prior and inject the speaker style embedding. After that, the proposed Flow-based User Emotion Controlling (FUEC) is used to synthesize waveform by flow matching prediction network conditioned on acoustics prior. In this process, the FUEC determines the gradient direction and guidance scale based on the user's emotion instructions by the positive and negative guidance mechanism, which focuses on amplifying the desired emotion while suppressing others. Extensive experimental results on three benchmark datasets demonstrate favorable performance compared to several state-of-the-art methods.

语音合成情感控制口型同步多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。