AI帮写幻灯片,更懂讲稿节奏与观众反应
DeepSlide: From Artifacts to Presentation Delivery

- 用分步时间预算规划讲稿逻辑链,控制每部分时长
- 生成的幻灯片与讲稿配合更紧密,节奏更精准
- 适合需要高质量演讲的科研、教学人员
演示文稿是学术传播的主要方式,但多数AI幻灯片生成器只关注视觉美观,忽视讲述节奏、叙事结构和准备过程。我们提出DeepSlide,一个面向完整演示流程的人机协同多智能体系统,覆盖需求分析、时间约束下的叙事规划、基于证据的幻灯片-讲稿生成、注意力增强及排练支持。DeepSlide集成:(i) 可控逻辑链规划器,为每个节点分配时间预算;(ii) 轻量级内容树检索器用于内容锚定;(iii) 带风格继承的马尔可夫式序列渲染;(iv) 沙箱执行与最小修复机制以保障可渲染性。我们还设计了双评分基准,明确区分静态幻灯片质量与动态演讲表现。在20个领域和多样观众群体中,DeepSlide在幻灯片质量上媲美强基线,但在交付指标上持续领先,显著提升叙事流畅度、节奏精确性和幻灯片-讲稿协同性,同时提供更清晰的关注引导。
原文摘要 · Abstract (English)
Presentations are a primary medium for scholarly communication, yet most AI slide generators optimize the artifact (a visually plausible deck) while under-optimizing the delivery process (pacing, narrative, and presentation preparation). We present DeepSlide, a human-in-the-loop multi-agent system that supports preparing the full presentation process, from requirement elicitation and time-budgeted narrative planning, to evidence-grounded slide--script generation, attention augmentation, and rehearsal support. DeepSlide integrates (i) a controllable logical-chain planner with per-node time budgets, (ii) a lightweight content-tree retriever for grounding, (iii) Markov-style sequential rendering with style inheritance, and (iv) sandboxed execution with minimal repair to ensure renderability. We further introduce a dual-scoreboard benchmark that cleanly separates static artifact quality from dynamic delivery excellence. Across 20 domains and diverse audience profiles, DeepSlide matches strong baselines on artifact quality while consistently achieving larger gains on delivery metrics, improving narrative flow, pacing precision, and slide--script synergy with clearer attention guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。