arXiv:2505.21497cs.CVcs.AI2025-05NeurIPS被引 60

将论文自动生成精美海报,精准传递核心内容。

Paper2Poster: Towards Multimodal Poster Automation from Scientific Papers

  • 分三步构建智能海报:解析论文、规划布局、渲染优化
  • 生成海报仅需0.005美元,比GPT-4o省87%算力
  • 支持中文开源模型,适合科研人员快速制作学术海报

学术海报生成是科学传播中的关键挑战,需将长篇交错的论文压缩为单页视觉连贯的展示。为此,我们首次提出海报生成的基准与评估体系,包含近年会议论文与其作者设计海报的配对数据,评估维度包括:(i) 视觉质量——与人类海报的语义一致性;(ii) 文本连贯性——语言流畅度;(iii) 整体评估——由视觉语言模型作为裁判,从六个细粒度维度评分;(iv) PaperQuiz——通过VLM答题测试海报传达论文核心内容的能力。基于此,我们提出PosterAgent,一种自上而下、视觉反馈驱动的多智能体系统:(a) 解析器提取论文结构化素材;(b) 规划器生成保留阅读顺序与空间平衡的二叉树布局;(c) 渲染-评论循环迭代优化每块面板,执行渲染代码并利用VLM反馈消除溢出、确保对齐。实验发现,尽管GPT-4o输出在视觉上吸引人,但存在文本噪声且PaperQuiz得分低;读者参与度是主要美学瓶颈,因人工海报高度依赖视觉语义表意。我们开源的Qwen-2.5系列模型在几乎所有指标上超越4o驱动系统,同时仅使用其87%的令牌。该系统可将22页论文转化为可编辑的.pptx格式海报,全程成本仅0.005美元。相关代码与数据集已公开于https://github.com/Paper2Poster/Paper2Poster。

原文摘要 · Abstract (English)

Academic poster generation is a crucial yet challenging task in scientific communication, requiring the compression of long-context interleaved documents into a single, visually coherent page. To address this challenge, we introduce the first benchmark and metric suite for poster generation, which pairs recent conference papers with author-designed posters and evaluates outputs on (i)Visual Quality-semantic alignment with human posters, (ii)Textual Coherence-language fluency, (iii)Holistic Assessment-six fine-grained aesthetic and informational criteria scored by a VLM-as-judge, and notably (iv)PaperQuiz-the poster's ability to convey core paper content as measured by VLMs answering generated quizzes. Building on this benchmark, we propose PosterAgent, a top-down, visual-in-the-loop multi-agent pipeline: the (a)Parser distills the paper into a structured asset library; the (b)Planner aligns text-visual pairs into a binary-tree layout that preserves reading order and spatial balance; and the (c)Painter-Commenter loop refines each panel by executing rendering code and using VLM feedback to eliminate overflow and ensure alignment. In our comprehensive evaluation, we find that GPT-4o outputs-though visually appealing at first glance-often exhibit noisy text and poor PaperQuiz scores, and we find that reader engagement is the primary aesthetic bottleneck, as human-designed posters rely largely on visual semantics to convey meaning. Our fully open-source variants (e.g. based on the Qwen-2.5 series) outperform existing 4o-driven multi-agent systems across nearly all metrics, while using 87% fewer tokens. It transforms a 22-page paper into a finalized yet editable .pptx poster - all for just $0.005. These findings chart clear directions for the next generation of fully automated poster-generation models. The code and datasets are available at https://github.com/Paper2Poster/Paper2Poster.

论文生成智能海报多模态自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。