用多智能体生成沉浸式有声绘本视频,故事更吸引人、画面角色一致。
MM-StoryAgent: Immersive Narrated Storybook Video Generation with a Multi-Agent Paradigm across Text, Image and Audio
- 设计多智能体框架,跨文本/图像/音频协同生成
- 多阶段写作提升故事吸引力,音画同步增强沉浸感
- 开源可替换模块,适合儿童内容创作者与研究者
大型语言模型(LLMs)与人工智能生成内容(AIGC)的快速发展推动了AI原生应用的发展,例如自动生成儿童喜爱的故事书。然而,故事吸引力不足、叙事表现力有限,且缺乏开源评估基准仍是挑战。为此,我们提出并开源了MM-StoryAgent,该系统能生成具有优化情节、角色一致图像和多通道音频的沉浸式有声绘本视频。其采用多智能体框架,利用LLMs及多种模态专家工具(生成模型与API)实现跨模态协同创作。通过多阶段写作流程提升故事质量,并融合音效、音乐与叙事内容增强沉浸体验。系统提供灵活的开源平台,支持生成模块替换。客观与主观评估验证了文本质量及跨模态对齐的有效性。演示与源码已公开。
原文摘要 · Abstract (English)
The rapid advancement of large language models (LLMs) and artificial intelligence-generated content (AIGC) has accelerated AI-native applications, such as AI-based storybooks that automate engaging story production for children. However, challenges remain in improving story attractiveness, enriching storytelling expressiveness, and developing open-source evaluation benchmarks and frameworks. Therefore, we propose and opensource MM-StoryAgent, which creates immersive narrated video storybooks with refined plots, role-consistent images, and multi-channel audio. MM-StoryAgent designs a multi-agent framework that employs LLMs and diverse expert tools (generative models and APIs) across several modalities to produce expressive storytelling videos. The framework enhances story attractiveness through a multi-stage writing pipeline. In addition, it improves the immersive storytelling experience by integrating sound effects with visual, music and narrative assets. MM-StoryAgent offers a flexible, open-source platform for further development, where generative modules can be substituted. Both objective and subjective evaluation regarding textual story quality and alignment between modalities validate the effectiveness of our proposed MM-StoryAgent system. The demo and source code are available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。