用多智能体精准编辑长篇幻灯片,错误率低于0.5%。
EditPPT: Faithful Long-Deck Slide Editing via Structured Tool-Using Multi-Agent with Dual-Modal Validators

- 通过本地工具调用限制指令空间,避免长文档错误蔓延。
- 在28个真实演讲稿上实现99.5%执行成功率,对象保留率达91.5%。
- 适合需要高保真长篇幻灯片自动编辑的办公与教育场景。
自动化幻灯片编辑需兼顾修改准确性、内容保真度及对长篇幅的鲁棒性。现有基于大模型的系统常因依赖理想化中间表示或开放式代码生成,在真实演示文件中表现不佳,易产生连锁错误。我们提出EditPPT,一种将幻灯片编辑重构为受限工具选择问题的多智能体框架。通过调用原生PowerPoint COM接口执行局部形状级操作,有效缩小大模型动作空间,同时保留用户原始结构。通过跨模态分离验证机制,双模态验证提升了指令遵循与视觉质量的评估可靠性。我们还构建了DeckEdit-Bench基准,包含28个真人撰写的演示文稿、582张幻灯片和183个编辑指令,覆盖短、中、长三类长度。实验表明,EditPPT整体执行率达99.5%,幻灯片定位F1为88.7%,指令遵循率82.5%,对象保留率91.5%,且在长篇文档上仍保持优异性能。代码与基准数据集已公开。
原文摘要 · Abstract (English)
Automating slide editing requires simultaneously satisfying modification accuracy, preservation fidelity, and robustness to deck length. Existing LLM-based systems often fail on real-world presentation files because they rely on idealized intermediate representations or open-ended code generation, which are prone to cascading errors in long decks. We introduce EditPPT, a multi-agent framework that reformulates slide editing as a constrained tool-selection problem. By executing localized shape-level operations through the native PowerPoint COM interface, EditPPT narrows the LLM action space while preserving the application-resolved structure of user-authored decks. By separating validation across modalities, our dual-modal validation provides more robust assessment of both instruction fidelity and visual quality. We also present DeckEdit-Bench, a benchmark with 28 human-authored decks, 582 slides, and 183 editing prompts across short, medium, and long deck tiers. Experiments show that EditPPT achieves a 99.5% execution rate, 88.7% slide-targeting F1, 82.5% instruction following, and 91.5% object preservation overall, while maintaining strong performance on long decks. Our code and benchmark are available at https://anonymous.4open.science/r/EditPPT-0E27/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。