arXiv:2509.15845eess.AS2025-09被引 3

一键生成角色分明、情绪丰富的多角色有声书。

Deep Dubbing: End-to-End Auto-Audiobook System with Text-to-Timbre and Context-Aware Instruct-TTS

  • 根据文本描述自动生成角色音色嵌入,实现音色定制。
  • 结合对话上下文与情感指令,合成富有表现力的语音。
  • 适合需要高效制作多角色有声书的创作者或出版方。

多角色有声书制作流程主要包括剧本分析、角色音色选择和语音合成三个阶段。其中剧本分析可通过自然语言处理模型高精度自动化,而角色音色选择仍依赖人工。语音合成通常采用人工配音或文本转语音(TTS)技术。尽管TTS提升了效率,但在情感表达、语调控制和场景适应性方面仍有不足。为此,我们提出DeepDubbing,一个端到端的多角色有声书自动化系统。该系统包含两个核心组件:文本到音色(Text-to-Timbre, TTT)模型与上下文感知指令式语音合成(Context-Aware Instruct-TTS, CA-Instruct-TTS)模型。TTT模型根据文本描述生成角色专属的音色嵌入;CA-Instruct-TTS模型通过分析对话上下文并融入细粒度情感指令,合成富有表现力的语音。该系统实现了角色音色匹配与情感丰富叙述的自动生成,为有声书制作提供了创新解决方案。

原文摘要 · Abstract (English)

The pipeline for multi-participant audiobook production primarily consists of three stages: script analysis, character voice timbre selection, and speech synthesis. Among these, script analysis can be automated with high accuracy using NLP models, whereas character voice timbre selection still relies on manual effort. Speech synthesis uses either manual dubbing or text-to-speech (TTS). While TTS boosts efficiency, it struggles with emotional expression, intonation control, and contextual scene adaptation. To address these challenges, we propose DeepDubbing, an end-to-end automated system for multi-participant audiobook production. The system comprises two main components: a Text-to-Timbre (TTT) model and a Context-Aware Instruct-TTS (CA-Instruct-TTS) model. The TTT model generates role-specific timbre embeddings conditioned on text descriptions. The CA-Instruct-TTS model synthesizes expressive speech by analyzing contextual dialogue and incorporating fine-grained emotional instructions. This system enables the automated generation of multi-participant audiobooks with both timbre-matched character voices and emotionally expressive narration, offering a novel solution for audiobook production.

有声书语音合成角色音色情感表达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。