arXiv:2512.04529cs.AI2025-12被引 24

让多个智能体协作生成既准确又美观的科研幻灯片。

SlideGen: Collaborative Multimodal Agents for Scientific Slide Generation

  • 用多个专用智能体分工完成故事架构、图文对齐和排版设计。
  • 在200篇论文上测试,布局更平衡,内容覆盖更全,文字更连贯。
  • 适合需要快速制作高质量学术报告的研究者使用。

从科学论文生成演示幻灯片不仅是段落摘要的问题,还需决定讲述什么故事、突出哪些图表与公式,并将其合理安排成视觉清晰而非拥挤重复的页面。同时处理长文本上下文与布局敏感设计,使论文转幻灯片成为一项独特的多模态挑战。现有方法多聚焦于文本内容选择,常导致幻灯片缺乏视觉平衡、叙事流畅性或多模态证据的连贯整合。本文提出SlideGen,一个协同的视觉-语言多智能体框架,协调叙事规划、多模态定位与布局构建。SlideGen分配专门智能体负责构建演示结构、将支持性图表与关键论点对齐、生成演讲备注,并通过多样化的布局库生成可编辑的PPTX文件。通过在整套幻灯片层面优化布局,系统生成既忠实于原始论文又适合作为演示的幻灯片集。为超越文本保真度评估,我们提出几何感知密度(GAD)指标,量化视觉杂乱、稀疏与碎片化,与人类判断高度一致。在包含四个互补维度的200篇论文基准上评估,SlideGen在布局平衡、内容覆盖率和文本连贯性方面均显著优于竞争基线。

原文摘要 · Abstract (English)

Creating presentation slides from scientific papers is not simply a matter of summarizing paragraphs. A presenter is required to decide what story to tell, which figures and equations to highlight, and how to arrange them into pages that are visually clear rather than crowded or repetitive. The need to jointly reason over long contexts and layout-sensitive design makes paper-to-slide generation a uniquely challenging multimodal task. Most existing approaches, however, focus mainly on textual content selection, producing slides that often lack visual balance, narrative flow, or coherent integration of multimodal evidence. In this work, we introduce SlideGen, a collaborative vision-language multi-agent framework that coordinates narrative planning, multimodal grounding, and layout composition. SlideGen assigns specialized agents to outline the presentation structure, align supporting figures and tables with key claims, generate speaker notes, and compose editable PPTX slides through a diverse layout library. By refining layouts at the deck level, the system produces slide decks that are both faithful to the source paper and effective as presentations. To evaluate slide generation beyond text fidelity, we propose geometry-aware density (GAD), a metric that captures visual clutter, sparsity, and fragmentation, and shows strong agreement with human judgments. Evaluated across four complementary dimensions on our 200-paper benchmark, SlideGen consistently and significantly improves layout balance, content coverage, and text coherence, outperforming competitive baselines in paper-to-slide generation. Our findings suggest that effective slide generation requires multimodal design reasoning, and that agent collaboration offers a principled bridge between document understanding and scientific communication.

幻灯片生成多智能体科研传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。