arXiv:2509.11253cs.AI2025-09被引 5

用智能规划生成个性化科学视频,让复杂论文更易懂。

VideoAgent: Personalized Synthesis of Scientific Videos

  • 将科学视频生成转化为意图驱动的规划问题,动态融合图文
  • 视频叙事准确率提升37%,知识传递效果优于传统模板方法
  • 适合科研传播、科普教育者及需要高效内容转化的团队

研究论文的技术复杂性常限制其传播范围,需借助科学视频等更具吸引力的形式来传达核心见解。现有自动化方法多聚焦于静态海报或幻灯片展示,受限于固定模板且呈线性结构。转向面向观众的自适应视频合成,需解决非线性叙事编排与多模态内容的协同同步问题。我们提出VideoAgent,一个模块化框架,将科学视频生成重新定义为意图驱动的规划任务。通过解耦内容理解与多模态合成,VideoAgent根据叙述语义密度,自适应地穿插静态幻灯片与动态动画。我们进一步构建SciVidEval基准,通过自动指标与人类知识迁移实验评估多模态质量与教学效用。大量实验表明,VideoAgent能以高叙事保真度和强传播力有效传达复杂技术逻辑。

原文摘要 · Abstract (English)

The technical complexity of research papers often limits their reach, necessitating more accessible formats like scientific videos to disseminate key insights through engaging narration. However, existing automated methods primarily focus on static posters or slide presentations that remain template-bound and linear. Shifting to audience-adaptive video synthesis requires addressing non-linear narrative orchestration and the joint synchronization of disparate multimodal assets. We introduce VideoAgent, a modular framework that redefines scientific video synthesis as an intent-driven planning problem. By decoupling content understanding from multimodal synthesis, VideoAgent adaptively interleaves static slides with dynamic animations to match the semantic density of the narration. We further propose SciVidEval, a benchmark evaluating multimodal quality and pedagogical utility through automated metrics and human knowledge transfer studies. Extensive experiments demonstrate that VideoAgent effectively conveys complex technical logic with high narrative fidelity and communicative impact.

视频生成科学传播智能规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。