arXiv:2602.03866cs.DLcs.AI2026-02被引 3

PaperX统一生成学术幻灯片,用结构化图谱提升质量与效率

PaperX: A Unified Framework for Multimodal Academic Presentation Generation with Scholar DAG

  • 构建学者DAG中间表示,解耦论文逻辑与展示语法
  • 多任务生成效果达当前最优,成本显著低于专用模型
  • 适合科研传播、自动化报告生成场景

将科学论文转化为多模态演示内容对研究传播至关重要,但目前仍高度依赖人工。现有自动化方案通常将每种格式视为独立下游任务,导致重复处理和语义不一致。我们提出PaperX,一个统一框架,将学术演示生成建模为结构转换与渲染过程。核心是学者DAG(Scholar DAG),一种中间表示,将论文的逻辑结构与最终呈现语法解耦。通过自适应图遍历策略,PaperX可从单一源生成多样且高质量输出。综合评估表明,该框架在内容保真度和美学质量上达到当前最优水平,同时相比专用单任务代理显著提升成本效益。

原文摘要 · Abstract (English)

Transforming scientific papers into multimodal presentation content is essential for research dissemination but remains labor intensive. Existing automated solutions typically treat each format as an isolated downstream task, leading to redundant processing and semantic inconsistency. We introduce PaperX, a unified framework that models academic presentation generation as a structural transformation and rendering process. Central to our approach is the Scholar DAG, an intermediate representation that decouples the paper's logical structure from its final presentation syntax. By applying adaptive graph traversal strategies, PaperX generates diverse, high quality outputs from a single source. Comprehensive evaluations demonstrate that our framework achieves the state of the art performance in content fidelity and aesthetic quality while significantly improving cost efficiency compared to specialized single task agents.

学术生成多模态结构化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。