通过逆向规划学习隐式设计意图,实现幻灯片级个性化生成
Personalization as Inverse Planning: Learning Latent Design Intents for Agentic Slide Generation via Structural Denoising

- 将幻灯片个性化建模为逆向规划问题,无需预设工具知识
- 利用结构去噪任务训练双智能体协作生成可执行设计,效果优于现有方法
- 适合需要精细排版定制的AI内容生成场景
幻灯片设计需同时个性化整体主题与页面布局。当前基于AI代理的方法在细粒度、页面级设计上表现不佳,仅依赖预设模板或用户冗长指令,难以捕捉隐式设计意图,导致页面级幻灯片个性化(PSP)问题未解。本文将PSP建模为逆向规划问题,提出无需预设具体执行工具(如PowerPoint、Beamer)的设计意图学习方法。为克服去端到端优化困难,提出SPIRE框架:通过人为破坏干净幻灯片的视觉结构,构建可验证的去噪任务,使两个智能体通过强化学习协同优化可执行设计。理论证明结构去噪是PSP的一致代理,且多智能体设定严格降低策略梯度方差。大量实验表明SPIRE显著优于现有方法。
原文摘要 · Abstract (English)
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or user verbose instructions, they fail to capture latent design intents, leaving Page-level Slide Personalization (PSP) unresolved. To close this gap, this work formulates PSP as an inverse planning problem. We propose to learn a design intent without assuming any knowledge of the specific executing tools (e.g., PowerPoint, Beamer) being used. However, relinquishing control over these tools makes the problem intractable to optimize end-to-end. To overcome this, we propose SPIRE, a principled framework to solve PSP approximately. By intentionally corrupting the visual structures of clean slides, SPIRE creates a verifiable task to denoise the corruption, whereby two agents learn to collaboratively refine executable designs via reinforcement learning (RL). We present a proof that structural denoising is a consistent surrogate for PSP, and that the multi-agent formulation strictly reduces policy gradient variance in RL. Extensive experiments demonstrate the superiority of SPIRE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。