arXiv:2601.23265cs.CLcs.CV2026-01被引 26

让AI自动生成符合发表标准的学术插图,解放科研人员精力。

PaperBanana: Automating Academic Illustration for AI Scientists

  • 用多个智能体协作完成参考检索、内容规划、图像渲染与自我优化。
  • 在292个神经网络会议插图测试中,各项指标均优于现有方法。
  • 不仅生成技术图,还能画出高质量统计图表,适合多领域研究者。

尽管基于语言模型的自主人工智能科学家发展迅速,但生成可发表的学术插图仍是研究流程中的繁重瓶颈。为解决这一问题,我们提出 PaperBanana,一个用于自动化生成可发表学术插图的代理框架。该框架利用先进的视觉语言模型(VLMs)和图像生成模型,通过多个专用智能体协同工作:检索参考、规划内容与风格、渲染图像,并通过自我批评实现迭代优化。为严格评估该框架,我们构建了 PaperBananaBench,包含从 NeurIPS 2025 论文中精选的 292 个方法学示意图测试案例,涵盖多样研究领域与绘图风格。综合实验表明,PaperBanana 在忠实度、简洁性、可读性和美学方面均显著优于主流基线方法。此外,我们的方法还可有效扩展至高质量统计图表生成。整体上,PaperBanana 为可发表插图的自动化生成铺平了道路。

原文摘要 · Abstract (English)

Despite rapid advances in autonomous AI scientists powered by language models, generating publication-ready illustrations remains a labor-intensive bottleneck in the research workflow. To lift this burden, we introduce PaperBanana, an agentic framework for automated generation of publication-ready academic illustrations. Powered by state-of-the-art VLMs and image generation models, PaperBanana orchestrates specialized agents to retrieve references, plan content and style, render images, and iteratively refine via self-critique. To rigorously evaluate our framework, we introduce PaperBananaBench, comprising 292 test cases for methodology diagrams curated from NeurIPS 2025 publications, covering diverse research domains and illustration styles. Comprehensive experiments demonstrate that PaperBanana consistently outperforms leading baselines in faithfulness, conciseness, readability, and aesthetics. We further show that our method effectively extends to the generation of high-quality statistical plots. Collectively, PaperBanana paves the way for the automated generation of publication-ready illustrations.

AI作图自动化科研效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。