零样本生成多角色有声书,自动匹配音色与情感
MultiActor-Audiobook: Zero-Shot Audiobook Generation with Faces and Voices of Multiple Speakers
- 用多模态角色画像和大模型指令生成,自动匹配角色音色与情绪
- 无需训练即可生成连贯且富有表现力的多角色朗读,效果媲美商用产品
- 适合需要快速制作多角色有声书的创作者或内容生产者
我们提出 MultiActor-Audiobook,一种零样本的有声书生成方法,可自动生成一致、富有表现力且符合角色特征的语调(包括语调与情感)。以往系统存在需手动配置语调、朗读单调或依赖昂贵训练等问题。我们的方法引入两项新流程:(1) MSP(多模态角色画像生成)与 (2) LSI(基于大语言模型的脚本指令生成),使系统在无额外训练情况下实现角色语调一致性与情感表达。通过人类评估与MLLM评估,结果表明其性能可与商用产品竞争。消融实验进一步验证了MSP与LSI的有效性。
原文摘要 · Abstract (English)
We introduce MultiActor-Audiobook, a zero-shot approach for generating audiobooks that automatically produces consistent, expressive, and speaker-appropriate prosody, including intonation and emotion. Previous audiobook systems have several limitations: they require users to manually configure the speaker's prosody, read each sentence with a monotonic tone compared to voice actors, or rely on costly training. However, our MultiActor-Audiobook addresses these issues by introducing two novel processes: (1) MSP (**Multimodal Speaker Persona Generation**) and (2) LSI (**LLM-based Script Instruction Generation**). With these two processes, MultiActor-Audiobook can generate more emotionally expressive audiobooks with a consistent speaker prosody without additional training. We compare our system with commercial products, through human and MLLM evaluations, achieving competitive results. Furthermore, we demonstrate the effectiveness of MSP and LSI through ablation studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。