arXiv:2605.30673cs.CL2026-05

构建首个多模态教学观察基准,助力AI评估课堂视频表现

TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation

论文配图:TeachObs: A Human-Validated Benchmark for Multimodal Teaching Observation and Model Evaluation
图 1 · 摘自论文原文
  • 基于30段跨国课堂视频,每15秒切片并人工标注39个教学行为
  • 五款前沿大模型在三类任务中表现不一,加中间帧反致误判增多
  • 既可细粒度分析行为,也支持整体教学评价,揭示AI与专家互补空间

课堂视频蕴含可观的教学实践信号,但其教育与视觉信息常缺乏适合模型评估的组织形式。本文提出 extit{TeachObs},一个面向多模态教学观察的人工验证基准。该数据集包含来自8个国家的30段公开授课视频,共划分为5,158个固定15秒片段。七位研究人员对每个片段标注了39个二值化观察标签,涵盖20个视觉类(如手势、板书、指向、教具使用)和19个非视觉类(如讲解、监控、提问、反馈、反思)。金标准标签采用基于克里彭多夫阿尔法的可靠性和普遍性规则构建。除片段级标注外,三位专家还对30节课进行课程级评分与质性评价,覆盖教学设计、教学实施、学习者反应、学习材料与课程收尾等方面。利用这两层人类标注参考,我们在三个任务上评估五款具备视觉能力的前沿大模型:纯文本片段编码、文本+图像帧片段编码、以及基于大模型为裁判的课程级覆盖评分。结果表明:无单一模型在所有任务中持续领先;添加中间帧会同时增加真阳性与假阳性;模型评估倾向于高估程序清晰的课程,相比专家判断。因此, extit{TeachObs} 支持细粒度标注基准与整课评估,揭示了人工智能在跨学科、多形态课堂中可辅助与需依赖专家判断的具体场景。

原文摘要 · Abstract (English)

Classroom videos contain observable teaching practices, but their pedagogical and visual signals are rarely organized in forms suitable for model evaluation. We present \textit{TeachObs}, a human-validated benchmark for multimodal teaching observation in classroom videos. \textit{TeachObs} includes 30 public lesson videos from eight countries divided into 5,158 fixed 15-second scenes. Seven researchers annotated each scene with 39 binary observation codes, covering 20 visual codes, such as gesture, board work, pointing, and visual materials, and 19 nonvisual codes, such as instruction, monitoring, questioning, feedback, and reflection. Gold segment labels are constructed using reliability- and prevalence-aware rules based on Krippendorff's alpha. In addition to segment-level labels, three expert raters produced lesson-level ratings and qualitative evaluations of instructional design, instructional delivery, learner response, learning materials, and lesson closure across the 30 lessons, with rater coverage detailed in the body. Using these two human reference layers, we evaluate five vision-capable frontier LLMs across three tracks - text-only segment coding, text + frame segment coding, and lesson-level coverage scored under an LLM-as-judge protocol - and find that no single model consistently outperforms others across all three tracks, that adding a mid-frame inflates both true and false attributions per scene, and that model evaluations over-rate procedurally clear lessons relative to expert raters. \textit{TeachObs} therefore supports both fine-grained annotation benchmarking and whole-lesson evaluation, showing where AI systems can assist classroom video analysis and where expert judgment remains necessary across varied subjects, classroom formats, and annotation difficulty levels.

多模态教学评估大模型评测课堂分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。