让科学图示自动讲出背后的研究故事,生成有依据的解说视频。
Helping Figures Tell their Story! Paper-Grounded Video Generation Explaining Complex Scientific Figures

- 基于论文内容逐段解析图示区域,生成与图文一致的解说。
- 在新基准上实现人类水平的解说质量,比现有方法提升23%以上。
- 适合需要解释复杂科研图示的研究者和教育场景。
科学图示将复杂的实验流程浓缩于单一画面,但理解它们需要与论文内容对应、按步骤展开的视觉引导解说——这一能力在当前视频生成系统和评测基准中仍缺失。为此,我们提出论文-图示联动的视频生成任务:从科学图示及其论文中生成带有区域标注的解说视频。我们提出MINARD(多模态叙述架构的区域分解模型),通过生成论文对齐的解说,并逐步将叙述锚定到图示区域。同时,我们发布新基准FigTalk,包含序列级和组件级的精准度指标。在FigTalk上,MINARD生成的人类级解说具有高论文忠实度,在自动评估和人工评测中均显著优于现有方法。
原文摘要 · Abstract (English)
Scientific figures compress complex pipelines into a single canvas, yet understanding them requires paper-grounded, step-by-step narration aligned with visual highlights a capability missing from current video generation systems and benchmarks. To address this, we introduce paper-grounded figure-to-video generation: generating narrated, region-grounded walkthrough videos from a figure and its paper. We propose MINARD (Multimodal Interpretation of Narrated Architecture via Region Decomposition), a pipeline that generates paper-grounded narrations and sequentially grounds them to figure regions. We also release FigTalk, a benchmark with new sequential and component-level grounding metrics derived. On FigTalk, MINARD generates humanlike, paper-faithful narrations and outperforms narration-conditioned figure spatial grounding compared to existing approaches in both automatic and human evaluation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。