arXiv:2605.09378cs.CVcs.AI2026-05被引 1

让科学教学视频连贯可信,自动保持知识一致性。

EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation

论文配图:EduStory: A Unified Framework for Pedagogically-Consistent Multi-Shot STEM Instructional Video Generation
图 1 · 摘自论文原文
  • 用教学状态追踪+脚本控制,让多镜头视频讲得通顺
  • 新评测基准显示叙事断裂减少40%以上,符合教学目标
  • 适合教育AI、视频生成研究者,尤其关注科学教学

长时序视频生成在画质上已取得进展,但现有方法在多镜头教学视频中仍难以维持知识一致性和连贯的教学叙事,尤其是在STEM领域。为此,我们提出EduStory,一个统一的可靠教学视频生成框架。该框架整合了教学状态建模以追踪持久的知识状态、脚本引导的结构化控制以组织多镜头叙事,以及面向学习的评估指标以衡量知识保真度与约束满足程度。为支持严格评估,我们进一步构建了EduVideoBench诊断基准,包含多粒度标注:教学故事板、镜头级语义和知识状态转移,并提供可控教学视频生成的基线任务。大量实验表明,领域感知的状态建模和结构化控制显著降低叙事断裂,提升与教学意图的对齐度。结果凸显了领域特定结构约束与定制化基准在推进可靠、可控、可信长时序视频生成中的重要性。

原文摘要 · Abstract (English)

Long-horizon video generation has advanced in visual quality, yet existing methods still struggle to maintain knowledge consistency and coherent pedagogical narratives across multi-shot instructional videos, especially in STEM domains. To address these challenges, we propose EduStory, a unified framework for reliable instructional video generation. EduStory integrates pedagogical state modeling to track persistent knowledge states, script-guided structured control to organize multi-shot narratives, and learning-oriented evaluation metrics to assess knowledge fidelity and constraint satisfaction. To support rigorous evaluation, we further introduce EduVideoBench, a diagnostic benchmark with multi-granularity annotations, including pedagogical storyboards, shot-level semantics, and knowledge state transitions, together with baseline tasks for controllable instructional video generation. Extensive experiments demonstrate that domain-aware state modeling and structured control substantially reduce narrative breakdown and improve alignment with instructional intent. These results highlight the significance of domain-specific structural constraints and tailored benchmarks for advancing reliable, controllable, and also trustworthy long-horizon video generation.

教学视频长时序生成STEM教育

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。