arXiv:2502.15412cs.CL2025-02被引 5

用自验证机制让AI自动生成更连贯美观的学术幻灯片

Textual-to-Visual Iterative Self-Verification for Slide Generation

  • 分内容与版式两步走,结合上下文提升信息相关性
  • 通过文本转视觉的自我审查流程,版式准确率显著提升
  • 适合需要高效制作学术汇报幻灯片的研究者使用

生成演示幻灯片是一项耗时任务,亟需自动化。现有基于大模型的自主代理因灵活性不足且缺乏自动优化机制,难以在实际中应用。我们将生成缺失幻灯片的任务分解为内容生成与版式生成两个关键环节,契合学术幻灯片制作流程。首先,提出一种内容生成方法,通过引入邻近幻灯片上下文并采用章节检索策略,增强内容连贯性与相关性。对于版式生成,提出基于大模型的审查-修正工作流(Reviewer + Refiner),实现从复杂文本布局到直观视觉格式的自验证转换。该模态转换简化任务,使审查与优化更贴近人工判断。实验表明,该方法在对齐度、逻辑流畅性、视觉吸引力和可读性方面显著优于基线方法。

原文摘要 · Abstract (English)

Generating presentation slides is a time-consuming task that urgently requires automation. Due to their limited flexibility and lack of automated refinement mechanisms, existing autonomous LLM-based agents face constraints in real-world applicability. We decompose the task of generating missing presentation slides into two key components: content generation and layout generation, aligning with the typical process of creating academic slides. First, we introduce a content generation approach that enhances coherence and relevance by incorporating context from surrounding slides and leveraging section retrieval strategies. For layout generation, we propose a textual-to-visual self-verification process using a LLM-based Reviewer + Refiner workflow, transforming complex textual layouts into intuitive visual formats. This modality transformation simplifies the task, enabling accurate and human-like review and refinement. Experiments show that our approach significantly outperforms baseline methods in terms of alignment, logical flow, visual appeal, and readability.

幻灯片生成大模型应用自验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。