arXiv:2506.23605cs.CVcs.AI2025-06中稿 · ICDAR 2025

用大模型生成逼真课件,解决标注数据少的问题

AI-Generated Lecture Slides for Improving Slide Element Detection and Retrieval

  • 用大模型自动生成高质量课件,替代人工标注
  • 仅用少量真实数据微调,性能显著优于纯真实数据训练
  • 适合需要课件理解但标注资源有限的研究者

课件元素检测与检索是课件理解的关键问题。现有方法依赖大量人工标注,但标注大量课件耗时且需领域知识。为此,我们提出基于大语言模型的合成课件生成流程 SynLecSlideGen,可生成高质量、连贯且逼真的课件。同时构建评估基准 RealSlide,手动标注了1,050张真实课件。通过在真实数据上进行少样本迁移学习验证合成数据有效性:使用合成数据预训练的模型,在仅少量真实数据下表现显著优于仅在真实数据上训练的模型。结果表明,合成数据能有效弥补标注数据不足。代码与资源已公开于项目网站:https://synslidegen.github.io/。

原文摘要 · Abstract (English)

Lecture slide element detection and retrieval are key problems in slide understanding. Training effective models for these tasks often depends on extensive manual annotation. However, annotating large volumes of lecture slides for supervised training is labor intensive and requires domain expertise. To address this, we propose a large language model (LLM)-guided synthetic lecture slide generation pipeline, SynLecSlideGen, which produces high-quality, coherent and realistic slides. We also create an evaluation benchmark, namely RealSlide by manually annotating 1,050 real lecture slides. To assess the utility of our synthetic slides, we perform few-shot transfer learning on real data using models pre-trained on them. Experimental results show that few-shot transfer learning with pretraining on synthetic slides significantly improves performance compared to training only on real data. This demonstrates that synthetic data can effectively compensate for limited labeled lecture slides. The code and resources of our work are publicly available on our project website: https://synslidegen.github.io/.

课件理解合成数据大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。