arXiv:2602.13318cs.AIcs.CV2026-02KDD被引 8

构建首个多智能体学术幻灯片生成与编辑评测基准

DECKBench: Benchmarking Multi-Agent Frameworks for Academic Slide Generation and Editing

  • 设计多轮编辑指令数据集,模拟真实教学场景
  • 从内容忠实度到布局质量全面评估幻灯片生成效果
  • 适合研究教育自动化、多智能体协作的学者参考

自动生成并迭代编辑学术幻灯片不仅需要文档摘要,还需忠实的内容选择、连贯的幻灯片组织、布局感知的渲染以及稳健的多轮指令遵循。然而现有评测基准和评估协议未能充分衡量这些挑战。为此,我们提出针对多智能体幻灯片生成与编辑的评测框架 DECKBench,基于精心筛选的论文-幻灯片配对数据集,并加入真实模拟的编辑指令。评估协议系统性地检验幻灯片级与整套幻灯片级的忠实度、连贯性、布局质量和多轮指令遵循能力。我们进一步实现了一个模块化多智能体基线系统,将任务分解为论文解析与摘要、幻灯片规划、HTML生成及迭代编辑。实验结果表明,该基准能凸显优势、暴露失效模式,并提供改进方向。本工作为学术演示文稿生成与编辑建立了可复现、可比较的标准化评估基础。代码与数据公开于 https://github.com/morgan-heisler/DeckBench。

原文摘要 · Abstract (English)

Automatically generating and iteratively editing academic slide decks requires more than document summarization. It demands faithful content selection, coherent slide organization, layout-aware rendering, and robust multi-turn instruction following. However, existing benchmarks and evaluation protocols do not adequately measure these challenges. To address this gap, we introduce the Deck Edits and Compliance Kit Benchmark (DECKBench), an evaluation framework for multi-agent slide generation and editing. DECKBench is built on a curated dataset of paper to slide pairs augmented with realistic, simulated editing instructions. Our evaluation protocol systematically assesses slide-level and deck-level fidelity, coherence, layout quality, and multi-turn instruction following. We further implement a modular multi-agent baseline system that decomposes the slide generation and editing task into paper parsing and summarization, slide planning, HTML creation, and iterative editing. Experimental results demonstrate that the proposed benchmark highlights strengths, exposes failure modes, and provides actionable insights for improving multi-agent slide generation and editing systems. Overall, this work establishes a standardized foundation for reproducible and comparable evaluation of academic presentation generation and editing. Code and data are publicly available at https://github.com/morgan-heisler/DeckBench .

幻灯片生成多智能体评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。