arXiv:2505.21660cs.LG2025-05EMNLP被引 15

用智能代理框架生成高质量图文并茂的演示文稿

PreGenie: An Agentic Framework for High-quality Visual Presentation Generation

  • 分两阶段迭代生成,利用多模态大模型协作优化排版与内容
  • 在美学和内容一致性上超越现有模型,更贴近人工设计偏好
  • 适合需要专业级演示文稿的科研、商务场景

视觉演示对有效沟通至关重要。早期基于深度学习的自动化生成方法常因布局混乱、文本摘要不准确及图像理解不足,导致图文不匹配,限制了其在商业和科研等正式场合的应用。为此,我们提出PreGenie,一个基于多模态大语言模型(MLLMs)的智能体式模块化框架,用于生成高质量视觉演示文稿。PreGenie基于Slidev框架,将幻灯片渲染为Markdown代码,分为两个阶段:(1) 分析与初始生成,总结多模态输入并生成初始代码;(2) 审查与重生成,迭代审查中间代码与渲染后的幻灯片以输出最终成果。每个阶段均调动多个MLLM协同工作并共享信息。全面实验表明,PreGenie在多模态理解能力上表现卓越,在美学与内容一致性方面优于现有模型,且更符合人类设计偏好。

原文摘要 · Abstract (English)

Visual presentations are vital for effective communication. Early attempts to automate their creation using deep learning often faced issues such as poorly organized layouts, inaccurate text summarization, and a lack of image understanding, leading to mismatched visuals and text. These limitations restrict their application in formal contexts like business and scientific research. To address these challenges, we propose PreGenie, an agentic and modular framework powered by multimodal large language models (MLLMs) for generating high-quality visual presentations. PreGenie is built on the Slidev presentation framework, where slides are rendered from Markdown code. It operates in two stages: (1) Analysis and Initial Generation, which summarizes multimodal input and generates initial code, and (2) Review and Re-generation, which iteratively reviews intermediate code and rendered slides to produce final, high-quality presentations. Each stage leverages multiple MLLMs that collaborate and share information. Comprehensive experiments demonstrate that PreGenie excels in multimodal understanding, outperforming existing models in both aesthetics and content consistency, while aligning more closely with human design preferences.

视觉生成多模态智能代理演示文稿

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。