让AI自动排版幻灯片,还能自我纠错,效果更好更省成本。
ACE: A Self-Correcting Agentic Canvas Editor for Multi-Slide Presentation Automation

- 用分层场景图+专用操作工具,精准控制幻灯片布局
- 自纠错机制让指令遵循度提升至4.23(基准3.81),速度快1.75倍
- 适合需要高质量自动化演示生成的团队或设计工具开发者
商业设计平台正通过大语言模型代理编辑文档,但面临两大难题:传统文档格式仅支持扁平化绝对定位元素,导致代理需重新计算坐标,常破坏布局;且设计无唯一标准答案,以差异对比为基准的评估指标会误判合理但不同的输出。我们提出ACE,一种基于分层场景图的智能画布编辑器,配备98种专用于演示文稿的操作工具,搭配内容感知路由模块CARE,可将输入令牌减少约89%;并引入无需真实标签的自纠错循环,由自然语言评价反馈作为下一回合指令。固定主干模型下,单次执行即达同类迭代式HTML流程性能;加入自纠错后,在94项任务基准上指令遵循得分提升至4.23(对比3.81,配对p=0.010),速度提升1.75倍,成本降低约44%。视觉质量无显著差异,但26名盲评者中,有58.7%明确偏好ACE,81%偏爱自纠错结果;该偏好在三类评判者中稳定一致,且外部评判者仍保留三分之二纠错收益,排除闭环依赖。66%案例一次完成,严格峰值回滚消除所有退化情况。
原文摘要 · Abstract (English)
Commercial design platforms increasingly edit documents through large language model (LLM) agents, but two practical problems block reliable deployment: legacy document formats expose only \emph{flat}, absolutely positioned elements, so agents must recompute coordinates and routinely break layouts; and design has no unique ground truth, so diff-against-reference metrics penalize valid-but-different outputs. We present \textbf{ACE}, an agentic canvas editor over a \emph{hierarchical scene-graph} with a presentation-specialized action space (98 tools), paired with \textbf{CARE}, a content-aware router that feeds the agent only the relevant slice of each deck (avg.\ $\sim$89\% input-token reduction), and a \emph{self-correction} loop driven by a \emph{ground-truth-free} instruction-following (IF) judge whose natural-language critique is fed back as the next-turn instruction. With a fixed backbone, a scene-graph editor in a \emph{single turn} already matches a same-backbone \emph{agentic} HTML pipeline that iterates internally; adding self-correction lifts ACE significantly above it on instruction following (IF 4.23 vs.\ 3.81 on the full 94-task benchmark, paired $p{=}.010$, replicated by an out-of-loop judge) at 1.75$\times$ the speed and $\sim$44\% lower cost. VQ means are statistically indistinguishable, but 26 blind raters prefer ACE overall (58.7\% decisive win-rate) and prefer the self-corrected output 81\% of the time; the ranking is invariant across three judge families, and out-of-loop judges retain two-thirds of the self-correction gain, bounding circularity. 66\% of cases halt after one pass, and a strict-peak rollback removes every observed regression.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。