测试五种视觉语言模型在零样本课堂参与度识别中的表现,发现效果差且对提示敏感。
Zero-Shot Vision-Language Models for Classroom Engagement Recognition: A Benchmark Study of Prompt Sensitivity and Cross-Dataset Generalization

- 用三种提示设计测试五种主流模型在两个教育数据集上的表现。
- 个体学生识别准确率接近随机,仅0.10的科恩卡帕系数;场景级识别可达0.60。
- 模型严重依赖提示词,同一图像准确率波动高达32个百分点,且部分模型被安全过滤拦截。
自动化课堂参与度识别在可扩展学习分析中具有巨大潜力,但现代视觉-语言模型(VLMs)在零样本条件下对此任务的适用性尚未深入探索。我们系统性地评估了五种广泛使用的VLMs:CLIP、BLIP-VQA、GPT-4o、LLaVA-1.5-7B和Qwen2.5VL-7B-Instruct,分别在两个互补的教育数据集上进行测试:DAiSEE(300个采样视频片段,个体学生视角)和学生课堂行为数据集(SCB,1,168张场景级图像)。每个模型采用三种提示变体——最小提示、基于评分标准的提示和思维链提示。实验揭示零样本VLM在参与度识别中的三大失败模式:(1)个体学生识别近乎随机,DAiSEE上科恩卡帕系数从未超过0.10;(2)严重类别坍缩,模型将85%-100%的预测分配给单一参与水平,与视觉内容无关;(3)极端提示敏感性,相同图像在不同提示下准确率波动达32个百分点。值得注意的是,场景级分类更具可行性:当使用行为锚定提示时,CLIP和GPT-4o的科恩卡帕系数约为0.60。此外,我们还发现部署障碍:GPT-4o的安全过滤器拒绝了98%涉及个体学生面部的思维链请求。研究结果为VLM在教育观察系统中的应用提供了基准,并揭示关键设计考量。
原文摘要 · Abstract (English)
Automated classroom engagement recognition holds substantial promise for scalable learning analytics, yet the suitability of modern Vision-Language Models (VLMs) for this task under zero-shot conditions remains largely unexplored. We present a systematic benchmark that evaluates five widely-used VLMs: CLIP, BLIP-VQA, GPT-4o, LLaVA-1.5-7B, and Qwen2.5VL-7B-Instruct across two complementary educational datasets: DAiSEE, an individual-student video dataset (300 sampled test clips), and the Student Classroom Behaviour dataset (SCB, 1,168 scene-level images). Each model is probed with three prompt variants spanning minimal, rubric-anchored, and chain-of-thought designs. Our experiments reveal three primary failure modes of zero-shot VLMs for engagement recognition: (1) near-random performance on individual students, with Cohen's kappa never exceeding 0.10 on DAiSEE; (2) severe class collapse, where models assign 85-100% of predictions to a single engagement level regardless of visual content; and (3) extreme prompt sensitivity, with accuracy swings of up to 32 percentage points on identical images depending solely on prompt phrasing. Remarkably, scene-level classification on SCB is substantially more tractable: CLIP and GPT-4o achieve kappa approximately 0.60 when prompted with behaviorally-grounded rubrics. We also document a practical barrier for deployment: GPT-4o's safety filters reject 98% of chain-of-thought requests involving individual student faces. Our findings provide a calibrated baseline and surface critical design considerations for the use of VLMs in educational observation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。