arXiv:2503.00455cs.SDcs.AI2025-03ACL被引 6

PodAgent让自动生成播客更自然,三步实现内容、配音与表达的统一。

PodAgent: A Comprehensive Framework for Podcast Generation

  • 用主持人-嘉宾-撰稿人多智能体协作生成深度讨论内容
  • 语音匹配准确率达87.4%,通过语音池实现角色适配
  • 大模型引导合成让对话更生动,适合内容创作与音视频开发

现有自动音频生成方法难以有效生成类播客音频节目。核心挑战在于深度内容生成与恰当富表现力的语音输出。本文提出PodAgent,一个完整的播客生成框架:1)设计主持人-嘉宾-撰稿人多智能体协作系统,生成信息丰富的话题讨论内容;2)构建语音池实现语音角色精准匹配;3)采用大语言模型增强的语音合成方法,生成富有表现力的对话语音。由于缺乏标准化评估标准,我们制定了全面的评估指南以有效衡量模型性能。实验结果表明,PodAgent在话题讨论对话生成上显著优于直接使用GPT-4,语音匹配准确率达87.4%,并通过大模型引导合成产生更具表现力的语音。演示页:https://podcast-agent.github.io/demo/。源码:https://github.com/yujxx/PodAgent。

原文摘要 · Abstract (English)

Existing Existing automatic audio generation methods struggle to generate podcast-like audio programs effectively. The key challenges lie in in-depth content generation, appropriate and expressive voice production. This paper proposed PodAgent, a comprehensive framework for creating audio programs. PodAgent 1) generates informative topic-discussion content by designing a Host-Guest-Writer multi-agent collaboration system, 2) builds a voice pool for suitable voice-role matching and 3) utilizes LLM-enhanced speech synthesis method to generate expressive conversational speech. Given the absence of standardized evaluation criteria for podcast-like audio generation, we developed comprehensive assessment guidelines to effectively evaluate the model's performance. Experimental results demonstrate PodAgent's effectiveness, significantly surpassing direct GPT-4 generation in topic-discussion dialogue content, achieving an 87.4% voice-matching accuracy, and producing more expressive speech through LLM-guided synthesis. Demo page: https://podcast-agent.github.io/demo/. Source code: https://github.com/yujxx/PodAgent.

播客生成多智能体语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。