arXiv:2601.09487cs.CL2026-01被引 7

构建可量化评估幻灯片生成的基准,提升评价客观性与一致性。

SlidesGen-Bench: Evaluating Slides Generation via Computational and Quantitative Metrics

  • 以视觉输出为评估基础,统一不同生成方法的评价框架。
  • 从内容、美学、可编辑性三方面提供可复现的量化指标。
  • 基于1500条人工偏好数据,显著提升与人类判断的一致性。

大型语言模型的快速发展催生了多样化的自动化幻灯片生成范式,涵盖代码驱动布局到图像中心合成。然而,现有评估协议难以在不同架构间提供可比分数,或依赖未经校准的人类判断。本文提出SlidesGen-Bench基准,遵循普适性、量化性与可靠性三大原则:首先,以视觉输出为分析对象,不依赖底层生成方式;其次,提出计算化评估方法,从内容、美学、可编辑性三个维度提供可复现的量化指标;最后,构建包含9种主流生成系统在7种场景下的1500条幻灯片的人工偏好对齐数据集(Slides-Align1.5k)。实验表明,该基准在与人类偏好对齐度上优于现有评估流程。代码与数据开源于https://github.com/YunqiaoYang/SlidesGen-Bench。

原文摘要 · Abstract (English)

The rapid evolution of Large Language Models (LLMs) has fostered diverse paradigms for automated slide generation, ranging from code-driven layouts to image-centric synthesis. However, evaluating these heterogeneous systems remains challenging, as existing protocols often struggle to provide comparable scores across architectures or rely on uncalibrated judgments. In this paper, we introduce SlidesGen-Bench, a benchmark designed to evaluate slide generation through a lens of three core principles: universality, quantification, and reliability. First, to establish a unified evaluation framework, we ground our analysis in the visual domain, treating terminal outputs as renderings to remain agnostic to the underlying generation method. Second, we propose a computational approach that quantitatively assesses slides across three distinct dimensions - Content, Aesthetics, and Editability - offering reproducible metrics where prior works relied on subjective or reference-dependent proxies. Finally, to ensure high correlation with human preference, we construct the Slides-Align1.5k dataset, a human preference aligned dataset covering slides from nine mainstream generation systems across seven scenarios. Our experiments demonstrate that SlidesGen-Bench achieves a higher degree of alignment with human judgment than existing evaluation pipelines. Our code and data are available at https://github.com/YunqiaoYang/SlidesGen-Bench.

幻灯片生成评估基准量化评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。