评测大模型对幻灯片布局的理解能力,发现现有模型重内容轻布局。
PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- 构建涵盖4类任务的幻灯片多模态评测集
- 模型能理解内容但无法合理排布元素,存在错位重叠
- 适合关注视觉结构推理与演示生成的研究者
幻灯片融合丰富文本与结构化视觉布局,是评估现代多模态大模型(MLLMs)多模态推理与布局理解能力的理想场景。然而,现有基准仅聚焦狭窄子任务,忽视以布局为中心的关键挑战,而这正是真实幻灯片创作与编辑的核心。为此,我们提出PPTBench,一个全面的多模态基准,用于评估大模型在幻灯片相关任务上的表现。基于958个PPTX文件,PPTBench涵盖4类任务共4,439个样本,包括检测、理解、修改与生成。实验揭示当前多模态大模型在语义理解与视觉布局推理间存在显著差距:模型可解析内容,但难以生成一致的空间排布。消融分析表明,现有模型难以将视觉线索与基于JSON的布局结构结合,也无法将视觉信息融入其API规划能力。案例研究直观暴露了系统性布局错误,如对齐偏差与元素重叠。这些发现为评估视觉-结构推理能力提供了新视角,并指明未来研究方向。所有数据集与代码均已开源,支持可复现性与后续研究。
原文摘要 · Abstract (English)
PowerPoint presentations combine rich textual content with structured visual layouts, making them a natural testbed for evaluating the multimodal reasoning and layout understanding abilities of modern MLLMs. However, existing benchmarks focus solely on narrow subtasks while overlooking layout-centric challenges, which are central to real-world slide creation and editing. To bridge this gap, we introduce PPTBench, a comprehensive multimodal benchmark for evaluating LLMs on PowerPoint-related tasks. Leveraging a diverse source of 958 PPTX files, PPTBench evaluates models across four categories with 4,439 samples, including Detection, Understanding, Modification, and Generation. Our experiments reveal a substantial gap between semantic understanding and visual-layout reasoning in current MLLMs: models can interpret slide content but fail to produce coherent spatial arrangements. Ablation and further analysis show that current MLLMs struggle to combine visual cues with JSON-based layout structures and fail to integrate visual information into their API planning ability. And case studies visually expose systematic layout errors such as misalignment and element overlap. These findings provides a new perspective on evaluating VLLMs in PPT scenarios, highlighting challenges and directions for future research on visual-structural reasoning and coherent slide generation. All datasets and code are fully released to support reproducibility and future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。