将长文档自动转为音画同步的讲解视频,像真人演示一样自然。
PresentAgent: Multimodal Agent for Presentation Video Generation

- 分步处理文档:切片、设计幻灯片、生成语音、对齐合成
- 在30组数据上接近人类水平,三项评估指标均表现优异
- 适合需要高效制作演示视频的研究者和教育工作者
我们提出PresentAgent,一个将长篇文档转化为带旁白讲解的演示视频的多模态智能体。现有方法仅能生成静态幻灯片或文字摘要,而PresentAgent突破限制,实现视觉与语音内容的精准同步,高度模拟真人演示效果。其模块化流程包括:系统性分割输入文档、规划并渲染幻灯片风格图像帧、利用大语言模型与文本转语音模型生成上下文相关语音,并实现音频与视觉的无缝对齐。针对多模态输出的复杂评估难题,我们构建了PresentEval——一个基于视觉-语言模型的统一评估框架,通过提示驱动的方式在内容保真度、视觉清晰度和观众理解度三个维度进行综合评分。在自建的30个文档-演示视频配对数据集上的实验表明,PresentAgent在所有评估指标上均接近人类水平。结果证明可控多模态智能体在将静态文本转化为动态、高效且可访问的演示形式方面具有巨大潜力。代码将于 https://github.com/AIGeeksGroup/PresentAgent 公开。
原文摘要 · Abstract (English)
We present PresentAgent, a multimodal agent that transforms long-form documents into narrated presentation videos. While existing approaches are limited to generating static slides or text summaries, our method advances beyond these limitations by producing fully synchronized visual and spoken content that closely mimics human-style presentations. To achieve this integration, PresentAgent employs a modular pipeline that systematically segments the input document, plans and renders slide-style visual frames, generates contextual spoken narration with large language models and Text-to-Speech models, and seamlessly composes the final video with precise audio-visual alignment. Given the complexity of evaluating such multimodal outputs, we introduce PresentEval, a unified assessment framework powered by Vision-Language Models that comprehensively scores videos across three critical dimensions: content fidelity, visual clarity, and audience comprehension through prompt-based evaluation. Our experimental validation on a curated dataset of 30 document-presentation pairs demonstrates that PresentAgent approaches human-level quality across all evaluation metrics. These results highlight the significant potential of controllable multimodal agents in transforming static textual materials into dynamic, effective, and accessible presentation formats. Code will be available at https://github.com/AIGeeksGroup/PresentAgent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。