让AI精准编辑幻灯片,保持格式与可编辑性。
SLIDEFORGE: An LLM Agent for Controllable Editing of Slides as Structured Artifacts

- 构建幻灯片状态图,融合视觉分解与PPT原生结构
- 在多种指标上优于现有方法,实现高质量重构
- 适合需要精细幻灯片修改的设计师和研究者
当前AI代理能出色描述幻灯片,但辅助编辑不仅需理解内容,还需保持版式、风格、组件结构和原生可编辑性。现有代理多基于截图或弱文档表示,常导致视觉单元破碎、可编辑内容栅格化或版式破坏。为此,我们提出可控制幻灯片编辑框架SLIDEFORGE,构建了幻灯片状态图(Deck State Graph),该图将视觉分解、原生pptx对象结构与感知组织相连接。通过恢复人类可识别组件并保留细粒度可编辑结构,SLIDEFORGE支持主题保持的幻灯片原生操作与渲染状态验证。我们还引入一种评估范式,联合衡量组件恢复、保留、重制一致性、视觉质量及原生可编辑性。实验表明,SLIDEFORGE在各项指标上均超越直接提示、基于截图的代理和通用代码代理基线。代码已公开于https://github.com/UIUC-MONET/SLIDEFORGE。
原文摘要 · Abstract (English)
Current AI agents compellingly describe slides. However, AI-assisted slide editing requires more than understanding: the output must retain layout, style, component structure, and native editability. Towards, AI-assisted slide editing, existing agents operate on screenshots or weak document representations and often fragment coherent visual units, rasterize editable content, or break layout. In contrast, for controllable slide editing, we introduce an agentic framework, SLIDEFORGE, which builds a Deck State Graph, an executable slide state that links visual decomposition, native pptx object structure, and perceptual organization. By recovering human-referable components while retaining fine-grained editable structure, SLIDEFORGE supports theme-preserving reconstruction through slide-native operations and rendered-state verification. We further introduce an evaluation paradigm for controllable slide transformation that jointly measures component recovery, preservation, restyling consistency, visual quality, and native editability. Experiments show that SLIDEFORGE outperforms direct prompting, screenshot-based agents, and generic code-agent baselines across these dimensions. Code is available at https://github.com/UIUC-MONET/SLIDEFORGE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。