arXiv:2602.18466cs.CYcs.AI2026-02被引 1

首个面向科学课视频的多模态评测基准,揭示大模型在教学推理上的局限。

Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos

  • 构建首个基于NGSS标准的科学课视频评测集SciIBI,含113段带教学实践标注的视频
  • 8个主流多模态大模型在区分相似教学行为时表现不佳,平均准确率不足60%
  • 模型常依赖表面线索而非真正理解教学逻辑,适合用于辅助专家评审而非替代

K-12科学课堂是学生通过话语协调现象、证据与解释模型的重要场所,但其多模态互动的复杂性使自动化分析困难。现有课堂话语评测主要聚焦数学且仅依赖文本转录,忽视了下一代科学教育标准(NGSS)强调的视觉素材与建模推理。为此,我们提出SciIBI——首个用于分析科学课堂话语的视频评测基准,包含113段符合NGSS标准的视频片段,标注了核心教学实践(CIP)及其复杂度等级。评估8个先进大模型和多模态大模型发现:当前模型难以区分语义相近的教学行为,表明CIP编码需要超越表层模式匹配的教学推理能力。此外,加入视频输入对不同架构带来不一致的性能提升。关键的是,基于证据的评估显示,模型常通过表面捷径成功,而非真正理解教学内涵。这些结果确立科学课堂话语为多模态AI的前沿挑战,并指向人机协作路径:模型应辅助专家快速检索证据,而非取代其判断。

原文摘要 · Abstract (English)

K-12 science classrooms are rich sites of inquiry where students coordinate phenomena, evidence, and explanatory models through discourse; yet, the multimodal complexity of these interactions has made automated analysis elusive. Existing benchmarks for classroom discourse focus primarily on mathematics and rely solely on transcripts, overlooking the visual artifacts and model-based reasoning emphasized by the Next Generation Science Standards (NGSS). We address this gap with SciIBI, the first video benchmark for analyzing science classroom discourse, featuring 113 NGSS-aligned clips annotated with Core Instructional Practices (CIP) and sophistication levels. By evaluating eight state-of-the-art LLMs and Multimodal LLMs, we reveal fundamental limitations: current models struggle to distinguish pedagogically similar practices, suggesting that CIP coding requires instructional reasoning beyond surface pattern matching. Furthermore, adding video input yields inconsistent gains across architectures. Crucially, our evidence-based evaluation reveals that models often succeed through surface shortcuts rather than genuine pedagogical understanding. These findings establish science classroom discourse as a challenging frontier for multimodal AI and point toward human-AI collaboration, where models retrieve evidence to accelerate expert review rather than replace it.

多模态教育AI评测基准教学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。