arXiv:2606.19256cs.AI2026-06

评测大模型生成幻灯片时对不同观众需求的适配能力。

X+Slides: Benchmarking Audience-Conditioned Slide Generation

论文配图:X+Slides: Benchmarking Audience-Conditioned Slide Generation
图 1 · 摘自论文原文
  • 基于8133个源文本探针,按观众类型分配权重动态评估。
  • 当前模型仅能覆盖71.4%的关键观众信息,仍有明显缺失。
  • 适合关注幻灯片生成真实性与观众适配性的研究者。

从源文档自动生成幻灯片是大语言模型的重要应用。现有基准主要评估幻灯片的完整性与技术深度,忽略了目标观众这一关键现实因素。例如,专家需要严谨证明,而决策者更关注可操作结论。为此,我们提出X+Slides,一个专为观众条件化幻灯片生成设计的基准。该基准基于涵盖113个主题和七种演示场景的多样化语料库,构建了由8133个去重、源文本对齐探针组成的动态评估框架。通过为同一源文本探针分配观众特定的效用权重,X+Slides报告四项互补指标:观众覆盖率衡量传达观众必要信息的程度,领域覆盖率显示各类信息的覆盖情况,效率度量单位注意力成本带来的实用价值,正确性验证幻灯片陈述是否得到源文本支持。在DeepPresenter、SlideTailor和NotebookLM上的实验表明,当前系统虽能恢复大量但不完整的观众必要信息:当τ_A=0.7时,DeepPresenter最佳观众覆盖率为0.714,SlideTailor为0.594,NotebookLM消融版本达0.853,但存在明显的接地差异。结果表明,视觉质量与广泛主题覆盖不能作为证据支持的依据,必须依赖源文本对齐评估。

原文摘要 · Abstract (English)

Automatically generating slide decks from source documents is an important application of large language models (LLMs). Existing benchmarks primarily assess slide completeness and technical depth, while overlooking the target audience as a critical real-world factor. For instance, specialists demand rigorous proofs, whereas decision-makers prioritize actionable conclusions. To bridge this gap, we introduce X+Slides, a benchmark specifically designed for audience-conditioned slide generation. Built on a diverse corpus spanning 113 topics and seven presentation scenes, X+Slides employs a dynamic evaluation framework constructed from 8,133 deduplicated, source-grounded probes. By assigning audience-specific utility weights to the same source-grounded probes, X+Slides reports four complementary metrics: Audience Coverage measures how much audience-essential information is conveyed, Domain-wise Coverage shows which information types are covered, Efficiency measures delivered utility per unit of attention cost, and Correctness verifies whether slide claims are supported by the source. Experiments on DeepPresenter, SlideTailor, and NotebookLM show that current systems can recover a substantial but still incomplete part of audience-essential information: at $τ_A=0.7$, DeepPresenter reaches a best Audience Coverage of 0.714, SlideTailor reaches 0.594, and the NotebookLM ablation reaches 0.853 while showing clear grounding differences. These results indicate that visual quality and broad topic coverage should not be treated as evidence support without source-grounded evaluation.

幻灯片生成观众适配评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。