arXiv:2505.04885cs.SDcs.HC2025-05被引 3

用AI生成带空间感的沉浸式有声书,让故事在耳边立体展开。

A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration

  • 多智能体框架协同生成角色音色与动态空间音效
  • 结合DTW和RNN实现剧情与音效精准同步
  • 适合教育、无障碍阅读及沉浸式叙事平台

本研究提出一种面向沉浸式有声书生产的多智能体AI框架。该框架采用FastSpeech 2与VALL-E进行神经文本转语音合成,实现富有表现力的叙述与角色专属音色;利用先进语言模型自动解析文本并生成逼真的空间音频效果。通过动态时间规整(DTW)与循环神经网络(RNN)实现音效与剧情的精确时序对齐。结合基于扩散模型的生成方法、高阶全向声学(HOA)与散射延迟网络(SDN),构建高度真实的三维声景,显著提升听众沉浸感与叙事真实感。该技术大幅拓展了有声书在教育内容、叙事平台及视障人群无障碍服务中的应用潜力。未来工作将聚焦个性化定制、合成语音的伦理管理,以及与多感官平台的融合。

原文摘要 · Abstract (English)

This research introduces an innovative AI-driven multi-agent framework specifically designed for creating immersive audiobooks. Leveraging neural text-to-speech synthesis with FastSpeech 2 and VALL-E for expressive narration and character-specific voices, the framework employs advanced language models to automatically interpret textual narratives and generate realistic spatial audio effects. These sound effects are dynamically synchronized with the storyline through sophisticated temporal integration methods, including Dynamic Time Warping (DTW) and recurrent neural networks (RNNs). Diffusion-based generative models combined with higher-order ambisonics (HOA) and scattering delay networks (SDN) enable highly realistic 3D soundscapes, substantially enhancing listener immersion and narrative realism. This technology significantly advances audiobook applications, providing richer experiences for educational content, storytelling platforms, and accessibility solutions for visually impaired audiences. Future work will address personalization, ethical management of synthesized voices, and integration with multi-sensory platforms.

有声书空间音频多智能体神经语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。