arXiv:2605.17356cs.CV2026-05

构建跨场景演示文稿生成统一评测基准,解决真实应用中评估不匹配问题。

UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings

论文配图:UniPPTBench: A Unified Benchmark for Presentation Generation Across Diverse Input Settings
图 1 · 摘自论文原文
  • 设计四类真实输入场景的统一评测集,覆盖模糊提示、长文档等复杂情况。
  • 提出场景感知评估协议,结合通用与专属指标,精准诊断系统能力短板。
  • 揭示当前模型在内容对齐、多源融合上的普遍缺陷,适合评估工具开发者使用。

现有研究多聚焦于单一输入场景下的演示文稿生成,而实际应用涵盖多种复杂情形,如模糊用户提示、长篇文档、多模态材料及多源异构数据。然而,当前评估方法缺乏场景针对性,主要依赖视觉美观、版式质量等通用标准,无法衡量不同输入场景所需的核心能力,如内容压缩的准确性、图文对齐性及跨源整合能力。为此,我们提出UniPPTBench,一个涵盖四种典型输入场景(模糊提示、长文档、多模态文档、多源生成)的统一基准,并引入场景感知的UniPPTEval评估协议,结合共享指标与场景特异性指标进行综合评价。同时提供透明基线以支持可复现比较。实验显示,不同场景下性能差异显著,且普遍存在内容锚定失败、多模态融合不足、跨源合成失效等问题。尤其发现,通用质量评分高的模型在需要内容准确性的场景中仍表现不佳。UniPPTBench与UniPPTEval共同为多样化现实场景下的演示文稿生成提供了可信、诊断性强的评估基础。代码与数据将公开。

原文摘要 · Abstract (English)

Existing works typically focus on presentation generation under isolated input settings, whereas real-world use cases span diverse scenarios, including vague user prompts, long documents, multimodal materials, and multiple heterogeneous sources. Moreover, current evaluations are often insufficiently scenario-specific. They mainly rely on generic presentation-quality criteria, such as visual appeal, layout quality, and overall coherence, but fail to assess the core capabilities required by different input settings, including grounded compression, visual-text alignment, and cross-source synthesis. Consequently, the field lacks a unified benchmark and a scenario-aware evaluation framework for faithfully diagnosing presentation-generation systems across diverse real-world settings. We present UniPPTBench, a unified benchmark for presentation generation across four representative input settings: vague-prompt, long-document, multimodal-document, and multi-source generation. We further introduce UniPPTEval, a scenario-aware evaluation protocol that combines shared metrics for cross-setting comparison with scenario-specific metrics tailored to the core requirements of each setting. We also provide transparent reference baselines to support reproducible comparison. Experiments on UniPPTBench reveal substantial performance variation across settings and recurring failure modes in content grounding, multimodal integration, and cross-source synthesis. In particular, strong performance on generic presentation-quality metrics does not necessarily imply strong task fulfillment in grounded scenarios. Together, UniPPTBench and UniPPTEval provide a faithful and diagnostic foundation for evaluating presentation generation across diverse real-world scenarios. Code and data will be publicly available.

演示生成评测基准多模态场景评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。