AI4Reading用多智能体协作自动生成中文有声书解读,更准确易懂。
AI4Reading: Chinese Audiobook Interpretation System Based on Multi-Agent Collaboration
- 11个专业智能体协同完成主题分析、案例提取与内容优化
- 生成脚本比人工更简洁准确,语音质量仍有提升空间
- 适合想快速获取书籍核心思想的读者和内容创作者
有声书解读日益受到关注,因其能为读者提供可访问且深入的书籍分析,带来实践洞察与思想启发。然而,其人工制作过程耗时耗力。为此,我们提出AI4Reading,一个基于大语言模型(LLMs)与语音合成技术的多智能体协作系统,用于生成类似播客的有声书解读。系统旨在实现三大目标:内容准确性、理解清晰度与逻辑叙事结构。我们构建了由11个专用智能体组成的框架,包括主题分析师、案例分析师、编辑、播音员和校对者,协同完成主题探索、现实案例提取、内容组织优化与自然口语化合成。对比专家解读与系统输出,结果表明尽管在语音生成质量上仍存差距,但生成的解读脚本更简洁且更准确。
原文摘要 · Abstract (English)
Audiobook interpretations are attracting increasing attention, as they provide accessible and in-depth analyses of books that offer readers practical insights and intellectual inspiration. However, their manual creation process remains time-consuming and resource-intensive. To address this challenge, we propose AI4Reading, a multi-agent collaboration system leveraging large language models (LLMs) and speech synthesis technology to generate podcast, like audiobook interpretations. The system is designed to meet three key objectives: accurate content preservation, enhanced comprehensibility, and a logical narrative structure. To achieve these goals, we develop a framework composed of 11 specialized agents,including topic analysts, case analysts, editors, a narrator, and proofreaders that work in concert to explore themes, extract real world cases, refine content organization, and synthesize natural spoken language. By comparing expert interpretations with our system's output, the results show that although AI4Reading still has a gap in speech generation quality, the generated interpretative scripts are simpler and more accurate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。