评测视觉语言模型对幻灯片的结构理解与抗干扰能力
VLM-SlideEval: Evaluating VLMs on Structured Comprehension and Perturbation Sensitivity in PPT
- 构建三轴评估框架:元素提取、扰动鲁棒性、叙事顺序恢复
- 模型在像素级提取上表现差,对文字/样式扰动有稳定响应
- 适合研究幻灯片理解或构建智能评阅系统的开发者
视觉语言模型(VLMs)被越来越多地用于评估多模态内容,包括演示文稿,但其针对幻灯片的理解仍缺乏深入研究。本文提出 VLM-SlideEval,一个从三个维度评估 VLMs 的框架:(1) 基于真实标注的元素级图像内容提取;(2) 对几何、风格和文本的可控扰动下的鲁棒性;(3) 从乱序幻灯片中恢复整套文稿叙事结构的高级理解能力。利用 Zenodo 公开数据集(https://huggingface.co/datasets/Forceless/Zenodo10K/viewer/default/pptx),我们标准化了 PowerPoint XML 与实时渲染结果中的元数据,形成统一可验证的结构。实证表明,当前 VLMs 在像素级提取任务中表现不佳,但在受控扰动下展现出非平凡的一致性、保真度与稳定性;单张幻灯片理解较好,但无法可靠捕捉跨幻灯片的叙事结构。这些结果揭示了现有 VLMs 在幻灯片评估中的局限性,推动开发具备校准能力的‘评论者内嵌’式评估器,以支持代理流程中的迭代优化与选择。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are increasingly used to evaluate multimodal content, including presentation slides, yet their slide-specific understanding remains underexplored {despite their growing role as critics in agentic, model-forward pipelines}. We introduce VLM-SlideEval, an evaluation framework that probes VLMs along three axes: (1) element-level extraction from slide images aligned to ground truth; (2) robustness to controlled perturbations in geometry, style, and text; and (3) higher-level comprehension, such as recovering a deck's narrative order from shuffled slides. Using publicly available decks from Zenodo (https://huggingface.co/datasets/Forceless/Zenodo10K/viewer/default/pptx), we standardize ground-truth element metadata from PowerPoint XML and live renderings into a unified, verifiable schema. Empirically, VLMs underperform on pixel-accurate extraction and show non-trivial agreement, fidelity, and consistency under controlled perturbations, while performing better on single-slide content understanding; however, they do not reliably capture narrative structure across slides. These results highlight the limits of current VLMs for slide evaluation and motivate calibrated, critic-in-the-loop evaluators that drive iterative refinement and selection in agentic pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。