无需训练的多智能体系统,让有声书更情感丰富、沉浸感强。
Dopamine Audiobook: A Training-free MLLM Agent for Emotional and Immersive Audiobook Generation
- 用大模型分角色设计语音与音效,实现精准对齐。
- 通过语调检索与自适应合成,生成自然富有情感的语音。
- 内置心理感知评估机制,更贴近人类听感偏好。
有声书生成旨在从多模态输入中创造丰富、沉浸的听觉体验,但现有方法面临三大挑战:(1) 多类音频(如语音、音效、音乐)缺乏协同生成及精确的时间与语义对齐;(2) 难以表达细腻情感,常导致机械式发音;(3) 缺乏与人类偏好一致的自动化评估框架。为此,我们提出 Dopamine Audiobook,一种统一的无训练多智能体系统,其中多模态大语言模型(MLLM)担任语音设计师和音频设计师双重角色,实现情感化、类人化、沉浸式有声书生成与评估。我们首先提出基于流的上下文感知框架,实现多类型音频在词级的语义与时间对齐;为增强表现力,引入词级副语言增强、句级语调检索与自适应TTS模型选择;最后,构建基于MLLM的评估框架,融合自我批判、换位思考与心理感知魔力情绪提示,确保评估结果既符合人类偏好又具任务迁移性。实验表明,该方法在多个指标上达到当前最优(SOTA)表现,且评估框架更贴近人类判断,并具备跨音频任务泛化能力。
原文摘要 · Abstract (English)
Audiobook generation aims to create rich, immersive listening experiences from multimodal inputs, but current approaches face three critical challenges: (1) the lack of synergistic generation of diverse audio types (e.g., speech, sound effects, and music) with precise temporal and semantic alignment; (2) the difficulty in conveying expressive, fine-grained emotions, which often results in machine-like vocal outputs; and (3) the absence of automated evaluation frameworks that align with human preferences for complex and diverse audio. To address these issues, we propose Dopamine Audiobook, a novel unified training-free multi-agent system, where a multimodal large language model (MLLM) serves two specialized roles (i.e., speech designer and audio designer) for emotional, human-like, and immersive audiobook generation and evaluation. Specifically, we firstly propose a flow-based, context-aware framework for diverse audio generation with word-level semantic and temporal alignment. To enhance expressiveness, we then design word-level paralinguistic augmentation, utterance-level prosody retrieval, and adaptive TTS model selection. Finally, for evaluation, we introduce a novel MLLM-based evaluation framework incorporating self-critique, perspective-taking, and psychological MagicEmo prompts to ensure human-aligned and self-aligned assessments. Experimental results demonstrate that our method achieves state-of-the-art (SOTA) performance on multiple metrics. Importantly, our evaluation framework shows better alignment with human preferences and transferability across audio tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。