arXiv:2605.26918cs.CL2026-05

首个教育视频生成评估基准,检验模型是否真正适合作为教学工具。

Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation

论文配图:Are Video Models Zero-Shot Learners and Reasoners in Education? EduVideoBench, A Knowledge-Skills-Attitude Benchmark for Educational Video Generation
图 1 · 摘自论文原文
  • 基于知识-技能-态度框架,系统评估教育视频的合理性与安全性。
  • 五款前沿视频模型在三方面均表现不足,距离课堂可用仍有差距。
  • 专家分析指出,节奏、字幕清晰度等细节错一处就可能让视频失效。

视频生成模型正快速进入课堂,但现有评估基准仅关注感知质量、内在真实性、通用安全性或视频作为推理媒介的能力,缺乏对输出内容教育有效性的评估。本文提出EduVideoBench,首个基于知识-技能-态度(KSA)框架的教育领域平衡性基准,使教学适切性与教育安全性得以联合评估,而非作为临时质量维度。在五款前沿视频生成模型上测试发现,其在知识、技能和态度三个维度均存在显著改进空间,尚不具备课堂部署条件。通过专家定性分析,我们进一步揭示教育有效性具有多成分特征:如节奏不当、字幕不清或符号使用错误等单一偏差,即可使原本正确的视频整体失效。我们期望EduVideoBench能引导视频生成模型向具备教学根基且课堂安全的方向发展。

原文摘要 · Abstract (English)

Video generation models (VGMs) are rapidly entering classrooms, yet existing benchmarks evaluate only perceptual quality, intrinsic faithfulness, generic safety, or video as a reasoning medium, and none assesses whether the outputs are educationally valid. In this work, we present EduVideoBench, the first balanced benchmark in the education domain, grounded in the Knowledge-Skills-Attitude (KSA) framework so that pedagogical adequacy and educational safety are evaluated jointly rather than as ad-hoc quality dimensions. Across five frontier VGMs, our results show substantial room for improvement across knowledge, skills, and attitude before they are classroom-ready. We complement this with a qualitative analysis of expert comments, finding that educational validity is multi-component, where a single misaligned element such as pacing, legibility, or notation can invalidate an otherwise correct video. We hope EduVideoBench will guide the development of VGMs that are pedagogically grounded and safe for the classroom.

教育视频生成评估KSA框架视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。