让幻灯片生成能像人一样边做边改,根据实际效果自动优化。
DeepPresenter: Environment-Grounded Reflection for Agentic Presentation Generation
- 用环境感知替代自我反思,实时检查幻灯片效果并修正
- 在多个场景下表现优于现有方法,90亿参数模型成本更低
- 适合需要高质量、可迭代生成幻灯片的用户
幻灯片生成需深入内容调研、连贯视觉设计及基于观察的迭代优化。但现有生成代理多依赖预设流程和固定模板。为此,我们提出 DeepPresenter,一个能适应多样用户意图、支持反馈驱动优化、超越脚本化流程的智能代理框架。DeepPresenter 自主规划、渲染并修订中间幻灯片成果,实现长周期迭代优化。更重要的是,它不依赖内部推理痕迹进行自我反思,而是基于感知到的幻灯片状态(如渲染结果)进行环境感知式反思,可在执行中识别并修正特定幻灯片问题。在覆盖多种生成场景的评估集上,DeepPresenter 达到当前最优性能,且微调后的 90 亿参数模型在显著更低成本下仍保持高竞争力。项目开源地址:https://github.com/icip-cas/PPTAgent
原文摘要 · Abstract (English)
Presentation generation requires deep content research, coherent visual design, and iterative refinement based on observation. However, existing presentation agents often rely on predefined workflows and fixed templates. To address this, we present DeepPresenter, an agentic framework that adapts to diverse user intents, enables effective feedback-driven refinement, and generalizes beyond a scripted pipeline. Specifically, DeepPresenter autonomously plans, renders, and revises intermediate slide artifacts to support long-horizon refinement with environmental observations. Furthermore, rather than relying on self-reflection over internal signals (e.g., reasoning traces), our environment-grounded reflection conditions the generation process on perceptual artifact states (e.g., rendered slides), enabling the system to identify and correct presentation-specific issues during execution. Results on the evaluation set covering diverse presentation-generation scenarios show that DeepPresenter achieves state-of-the-art performance, and the fine-tuned 9B model remains highly competitive at substantially lower cost. Our project is available at: https://github.com/icip-cas/PPTAgent
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。