评测智能体在制作与编辑幻灯片上的表现,更真实地反映其实际能力。
PPT-Eval: A Benchmark for Computer-Use Agents on PowerPoint Tasks

- 设计分层级任务与评分规则,支持部分完成的打分。
- 人类评价相关性达0.77,验证评估框架有效性。
- 现有顶尖模型仅45%完全成功,适合评估办公自动化智能体。
制作和编辑幻灯片是专业与教育场景中常见的多模态活动,非常适合测试真实世界的计算机使用智能体。微软PowerPoint是应用最广泛、功能最丰富的演示文稿创建环境之一。我们提出PPT-Eval,一个包含120个任务的基准测试,覆盖12个文件,涵盖内容创作与演示编辑两类场景,并按难度分级。该领域核心挑战在于评估:任务复杂、多模态且常有多种合理解法。当前智能体常仅完成部分工作,传统二元成功指标无法捕捉进展。为此,我们设计了鲁棒的评估框架,基于过往评分体系构建任务专属评分标准,对中间步骤给予部分得分,惩罚多余修改与不良美学表现,并提供自然语言反馈。该方法与人工评判的相关性达到Kendall's τ-b 0.77。我们发现,现有前沿智能体仍难以有效解决幻灯片任务,如Claude-4.5-Opus仅达45%成功率,平均部分得分57%。基准数据集可访问:https://microsoft.github.io/ppteval。
原文摘要 · Abstract (English)
Creating and editing slides is a rich, multimodal activity that is ubiquitous in professional and educational settings, making it an ideal testbed for real-world computer-use agents. Microsoft PowerPoint is among the most widely adopted and feature-rich environments for presentation creation. We introduce PPT-Eval, a benchmark of 120 PowerPoint tasks across 12 files that cover both content creation and presentation editing scenarios, organized by difficulty. A central challenge in this domain is evaluation: tasks are complex, multimodal, and often admit many valid solutions. Moreover, today's agents frequently make only partial progress, which binary success metrics fail to capture. To address this, we design a robust evaluation framework to help create task-specific rubrics for PowerPoint tasks, taking inspiration from and building on past works for rubric-based evaluation. These rubrics award partial credit for intermediate steps, penalize unnecessary changes and poor aesthetics, and provide natural language feedback. This nuanced approach proves highly effective, achieving a Kendall's τ-b correlation of 0.77 with human judgments. We find that existing frontier agents still struggle with solving PowerPoint tasks, with strong models like Claude-4.5-Opus achieving only a 45% success rate and an average partial score of 57%. The benchmark is located at: https://microsoft.github.io/ppteval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。