无需真实参考稿,用负样本让大模型自动评估幻灯片质量并给出改进建议。
Taming LLMs with Negative Samples: A Reference-Free Framework to Evaluate Presentation Content with Actionable Feedback
- 通过生成带特定缺陷的负样本幻灯片,训练大模型理解内容评价标准。
- 在多个指标上优于传统方法和现有大模型评估,评分与反馈更准确。
- 适合需要自动化幻灯片质量评估的研究者或教育技术开发者。
自动生成演示文稿是生成式AI时代的重要课题。本文聚焦于评估幻灯片中多模态内容的有效性,即能否准确总结文档并面向广泛受众传达概念。我们构建了基准数据集RefSlides,包含多种主题的人工制作高质量演示文稿。接着提出一组用于刻画演示文稿内在特性的度量指标,并设计REFLEX评估框架,通过生成带有不同度量特异性扰动的负样本幻灯片来微调大语言模型,实现无参考的评分与可操作反馈。该方法在推理阶段无需真实参考演示文稿。大量自动与人工实验表明,本方法在生成分数与解释方面均显著优于经典启发式方法及当前最先进的大语言模型评估方法。
原文摘要 · Abstract (English)
The generation of presentation slides automatically is an important problem in the era of generative AI. This paper focuses on evaluating multimodal content in presentation slides that can effectively summarize a document and convey concepts to a broad audience. We introduce a benchmark dataset, RefSlides, consisting of human-made high-quality presentations that span various topics. Next, we propose a set of metrics to characterize different intrinsic properties of the content of a presentation and present REFLEX, an evaluation approach that generates scores and actionable feedback for these metrics. We achieve this by generating negative presentation samples with different degrees of metric-specific perturbations and use them to fine-tune LLMs. This reference-free evaluation technique does not require ground truth presentations during inference. Our extensive automated and human experiments demonstrate that our evaluation approach outperforms classical heuristic-based and state-of-the-art large language model-based evaluations in generating scores and explanations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。