arXiv:2508.09632cs.CVcs.AI2025-08ICCV被引 8

将论文自动转为结构化视频摘要,支持多领域高质量生成。

Preacher: Paper-to-Video Agentic System

  • 分层式流程:先分解论文,再逐段生成视频片段。
  • 跨模态对齐提升效果,生成视频在5个领域表现优异。
  • 适合科研传播、学术展示,提升论文影响力。

论文到视频任务旨在将研究论文转化为结构化视频摘要,提炼关键概念、方法与结论,使其更易理解且条理清晰。尽管当前顶尖视频生成模型展现潜力,但仍受限于上下文窗口小、视频时长固定、风格单一以及无法表达领域知识。为此,我们提出Preacher——首个论文到视频的智能体系统。Preacher采用自上而下的策略,对论文进行分解、总结与重构,再通过自下而上的方式生成视频,将多样化的视频片段合成连贯的摘要。为对齐跨模态表征,我们定义关键场景,并引入渐进式思维链(P-CoT)实现细粒度、迭代式规划。Preacher在五个研究领域成功生成高质量视频摘要,展现出超越现有视频生成模型的专业能力。代码将公开于:https://github.com/Gen-Verse/Paper2Video

原文摘要 · Abstract (English)

The paper-to-video task converts a research paper into a structured video abstract, distilling key concepts, methods, and conclusions into an accessible, well-organized format. While state-of-the-art video generation models demonstrate potential, they are constrained by limited context windows, rigid video duration constraints, limited stylistic diversity, and an inability to represent domain-specific knowledge. To address these limitations, we introduce Preacher, the first paper-to-video agentic system. Preacher employs a topdown approach to decompose, summarize, and reformulate the paper, followed by bottom-up video generation, synthesizing diverse video segments into a coherent abstract. To align cross-modal representations, we define key scenes and introduce a Progressive Chain of Thought (P-CoT) for granular, iterative planning. Preacher successfully generates high-quality video abstracts across five research fields, demonstrating expertise beyond current video generation models. Code will be released at: https://github.com/Gen-Verse/Paper2Video

论文生成视频摘要智能体系统跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。