arXiv:2510.05096cs.CVcs.AI2025-10被引 27

让论文自动变成学术演讲视频,省去大量人工制作时间。

Paper2Video: Automatic Video Generation from Scientific Papers

  • 用多智能体框架自动生成幻灯片、字幕、语音和虚拟讲者
  • 在101篇论文数据集上生成视频,信息保留率显著优于基线
  • 适合需要快速传播研究成果的科研人员和学术机构

学术演讲视频已成为科研传播的重要媒介,但制作过程高度依赖人工,通常需数小时设计幻灯片、录制与剪辑才能完成2至10分钟的视频。与自然视频不同,学术视频生成面临独特挑战:输入来自科研论文,包含密集的多模态信息(文本、图表、表格),且需协调幻灯片、字幕、语音与虚拟讲者等多通道对齐。为此,我们提出Paper2Video,首个包含101篇论文及其作者制作的演讲视频、幻灯片与讲者元数据的基准数据集。我们设计了四项定制评估指标——元相似性、PresentArena、PresentQuiz和IP Memory,用于衡量视频向观众传达论文信息的能力。在此基础上,我们提出PaperTalker,首个面向学术演讲视频生成的多智能体框架。它通过创新的树搜索视觉选择策略实现幻灯片布局优化,结合光标定位、字幕生成、语音合成与虚拟人渲染,并行化每页幻灯片生成以提升效率。在Paper2Video上的实验表明,所生成视频比现有基线更忠实、信息量更高,为自动化、可直接使用的学术视频生成迈出了切实一步。数据集、模型与代码已开源:https://github.com/showlab/Paper2Video。

原文摘要 · Abstract (English)

Academic presentation videos have become an essential medium for research communication, yet producing them remains highly labor-intensive, often requiring hours of slide design, recording, and editing for a short 2 to 10 minutes video. Unlike natural video, presentation video generation involves distinctive challenges: inputs from research papers, dense multi-modal information (text, figures, tables), and the need to coordinate multiple aligned channels such as slides, subtitles, speech, and human talker. To address these challenges, we introduce Paper2Video, the first benchmark of 101 research papers paired with author-created presentation videos, slides, and speaker metadata. We further design four tailored evaluation metrics--Meta Similarity, PresentArena, PresentQuiz, and IP Memory--to measure how videos convey the paper's information to the audience. Building on this foundation, we propose PaperTalker, the first multi-agent framework for academic presentation video generation. It integrates slide generation with effective layout refinement by a novel effective tree search visual choice, cursor grounding, subtitling, speech synthesis, and talking-head rendering, while parallelizing slide-wise generation for efficiency. Experiments on Paper2Video demonstrate that the presentation videos produced by our approach are more faithful and informative than existing baselines, establishing a practical step toward automated and ready-to-use academic video generation. Our dataset, agent, and code are available at https://github.com/showlab/Paper2Video.

视频生成多模态学术传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。