arXiv:2601.00943cs.CV2026-01中稿 · IEEE/CVF Winter Co…被引 3

构建物理教育视频生成基准,评估AI生成视频的科学准确性。

PhyEduVideo: A Benchmark for Evaluating Text-to-Video Models for Physics Education

  • 将物理概念拆解为教学点,设计精准提示词生成可视化讲解。
  • 当前模型视频流畅但概念准确率低,电磁学与热力学表现最差。
  • 适合教育AI研究者、课程开发人员使用,推动个性化学习发展。

生成式AI,尤其是文本到视频(T2V)系统,为科学教育提供了自动化生成生动直观视觉解释的前景。本文首次针对物理教育场景,提出一个专门用于评估解释性视频生成的基准。该基准通过将每个物理概念分解为细粒度教学点,并为每个点设计精心构造的提示词,以评估T2V模型生成视觉解释的能力。我们旨在系统探索利用T2V模型生成高质量、符合课程标准教育内容的可行性,推动可扩展、可访问、个性化的智能学习体验。评估结果表明,当前模型能生成视觉连贯、运动平滑且闪烁极少的视频,但在概念准确性方面可靠性不足。力学、流体和光学领域表现尚可,而电磁学与热力学等抽象交互领域则表现不佳。这凸显了视觉质量与概念正确性之间的差距。我们希望该基准能助力社区缩小这一差距,迈向可规模化生成准确、符合课程要求的物理教育视频的AI系统。基准与代码库已公开于 https://github.com/meghamariamkm/PhyEduVideo。

原文摘要 · Abstract (English)

Generative AI models, particularly Text-to-Video (T2V) systems, offer a promising avenue for transforming science education by automating the creation of engaging and intuitive visual explanations. In this work, we take a first step toward evaluating their potential in physics education by introducing a dedicated benchmark for explanatory video generation. The benchmark is designed to assess how well T2V models can convey core physics concepts through visual illustrations. Each physics concept in our benchmark is decomposed into granular teaching points, with each point accompanied by a carefully crafted prompt intended for visual explanation of the teaching point. T2V models are evaluated on their ability to generate accurate videos in response to these prompts. Our aim is to systematically explore the feasibility of using T2V models to generate high-quality, curriculum-aligned educational content-paving the way toward scalable, accessible, and personalized learning experiences powered by AI. Our evaluation reveals that current models produce visually coherent videos with smooth motion and minimal flickering, yet their conceptual accuracy is less reliable. Performance in areas such as mechanics, fluids, and optics is encouraging, but models struggle with electromagnetism and thermodynamics, where abstract interactions are harder to depict. These findings underscore the gap between visual quality and conceptual correctness in educational video generation. We hope this benchmark helps the community close that gap and move toward T2V systems that can deliver accurate, curriculum-aligned physics content at scale. The benchmark and accompanying codebase are publicly available at https://github.com/meghamariamkm/PhyEduVideo.

视频生成教育AI物理教育评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。