arXiv:2605.11363cs.CVcs.CL2026-05被引 2

让AI根据用户提问自动生成带多模态内容和互动的讲解视频。

PresentAgent-2: Towards Generalist Multimodal Presentation Agents

论文配图:PresentAgent-2: Towards Generalist Multimodal Presentation Agents
图 1 · 摘自论文原文
  • 基于用户问题自动检索资料,生成包含图文动效的完整视频。
  • 支持单人讲解、多人讨论和观众问答三种模式,可交互。
  • 首次构建多场景评测基准,覆盖内容质量与对话自然度等维度。

演示生成正从静态幻灯片迈向端到端的视频生成,融合研究依据、多模态媒体与交互式呈现。我们提出 PresentAgent-2,一个从用户查询生成演示视频的智能体框架。给定开放式问题与选定的演示模式后,PresentAgent-2 首先将问题提炼为聚焦主题,并在适配演示的资源中进行深度调研,收集文本、图像、GIF 和视频等多模态素材。随后构建幻灯片、生成特定模式脚本,并将幻灯片、音频与动态媒体合成完整演示视频。该框架支持三种独立模式:单人讲解(Single Presentation),生成单讲者叙述视频;讨论模式(Discussion),创建具结构化角色的多讲者演示,如提问引导、概念解释、细节澄清与要点总结;交互模式(Interaction),可基于生成的幻灯片、脚本、检索证据与上下文回答观众提问。为评估这些能力,我们构建了一个涵盖单人讲解、讨论与交互场景的多模态演示基准,设有任务特定的评价标准,包括内容质量、媒体相关性、动态媒体使用、对话自然度与交互准确性。总体而言,PresentAgent-2 将演示生成从依赖文档的幻灯片制作,拓展至以查询驱动、研究支撑、多模态媒体、对话与交互为核心的视频生成。代码:https://github.com/AIGeeksGroup/PresentAgent-2。网站:https://aigeeksgroup.github.io/PresentAgent-2。

原文摘要 · Abstract (English)

Presentation generation is moving beyond static slide creation toward end-to-end presentation video generation with research grounding, multimodal media, and interactive delivery. We introduce PresentAgent-2, an agentic framework for generating presentation videos from user queries. Given an open-ended user query and a selected presentation mode, PresentAgent-2 first summarizes the query into a focused topic and performs deep research over presentation-friendly sources to collect multimodal resources, including relevant text, images, GIFs, and videos. It then constructs presentation slides, generates mode-specific scripts, and composes slides, audio, and dynamic media into a complete presentation video. PresentAgent-2 supports three independent presentation modes within a unified framework: Single Presentation, which generates a single-speaker narrated presentation video; Discussion, which creates a multi-speaker presentation with structured speaker roles, such as for asking guiding questions, explaining concepts, clarifying details, and summarizing key points; and Interaction, which independently supports answering audience questions grounded in the generated slides, scripts, retrieved evidence, and presentation context. To evaluate these capabilities, we build a multimodal presentation benchmark covering single presentation, discussion, and interaction scenarios, with task-specific evaluation criteria for content quality, media relevance, dynamic media use, dialogue naturalness, and interaction grounding. Overall, PresentAgent-2 extends presentation generation from document-dependent slide creation to query-driven, research-grounded presentation video generation with multimodal media, dialogue, and interaction. Code: https://github.com/AIGeeksGroup/PresentAgent-2. Website: https://aigeeksgroup.github.io/PresentAgent-2.

演示生成多模态交互式智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。