arXiv:2601.14777cs.CVcs.AI2026-01被引 3

构建首个中文影视配音数据集并推出通用配音模型,支持多场景零样本配音。

FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

  • 自动生成大规模高质量中文影视配音数据,含丰富标注。
  • 在对话、多人场景中表现优于现有方法,语音质量与口型同步更精准。
  • 适合影视自动化配音、跨语言内容本地化等应用开发。

影视配音需根据视频场景合成语音,要求口型同步精准、音色还原真实,并准确表达角色身份与情感。然而现有方法存在两大瓶颈:(1) 高质量多模态配音数据规模小、词错误率高、标注稀疏、依赖人工标注,且仅限独白场景,制约模型训练;(2) 现有模型仅依赖唇部区域学习视听对齐,在复杂实景电影场景中适用性差,口型同步、语音质量与情感表达均不理想。为此,我们提出 FunCineForge,包含端到端的大规模配音数据生成流程和面向多样电影场景的基于多模态大模型的配音模型。利用该流程,我们构建了首个富含标注的中文电视剧配音数据集,并验证了其高质量。在独白、旁白、对话及多人场景上的实验表明,该模型在语音质量、口型同步、音色迁移和指令遵循方面持续超越当前最优方法。代码与演示见 https://anonymous.4open.science/w/FunCineForge。

原文摘要 · Abstract (English)

Movie dubbing is the task of synthesizing speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. However, existing methods face two major limitations: (1) high-quality multimodal dubbing datasets are limited in scale, suffer from high word error rates, contain sparse annotations, rely on costly manual labeling, and are restricted to monologue scenes, all of which hinder effective model training; (2) existing dubbing models rely solely on the lip region to learn audio-visual alignment, which limits their applicability to complex live-action cinematic scenes, and exhibit suboptimal performance in lip sync, speech quality, and emotional expressiveness. To address these issues, we propose FunCineForge, which comprises an end-to-end production pipeline for large-scale dubbing datasets and an MLLM-based dubbing model designed for diverse cinematic scenes. Using the pipeline, we construct the first Chinese television dubbing dataset with rich annotations, and demonstrate the high quality of these data. Experiments across monologue, narration, dialogue, and multi-speaker scenes show that our dubbing model consistently outperforms SOTA methods in audio quality, lip sync, timbre transfer, and instruction following. Code and demos are available at https://anonymous.4open.science/w/FunCineForge.

影视配音多模态语音合成数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。