arXiv:2506.10334cs.CVcs.AI2025-06被引 4

用视觉语言模型分析学生面部表情,零样本识别学习情绪。

Using Vision Language Models to Detect Students' Academic Emotion through Facial Expressions

  • 直接用大模型零样本识别面部情绪,无需额外训练。
  • 模型能准确识别开心,但难发现分心状态。
  • 通义千问表现更好,适合检测学生困惑场景。

学生的学习情绪显著影响其社交行为与学习效果。传统自动分析方法依赖监督学习,但泛化能力差,需反复收集标注数据并训练。视觉语言模型(VLMs)通过零样本提示实现跨任务泛化,无需微调。本研究探索VLM在在线学习环境中通过面部表情分析学生学术情绪的潜力。使用Llama-3.2-11B-Vision-Instruct和Qwen2.5-VL-7B-Instruct两个模型,对5000张包含困惑、分心、开心、中性、疲惫表情的图像进行零样本分析。初步结果显示,两模型均表现出中等识别性能,其中Qwen2.5-VL-7B-Instruct优于前者。两者在识别开心情绪上表现优异,但难以检测分心行为;而Qwen2.5-VL-7B-Instruct在识别困惑表情方面表现相对较好,显示出在识别引发学生困惑内容方面的应用潜力。

原文摘要 · Abstract (English)

Students' academic emotions significantly influence their social behavior and learning performance. Traditional approaches to automatically and accurately analyze these emotions have predominantly relied on supervised machine learning algorithms. However, these models often struggle to generalize across different contexts, necessitating repeated cycles of data collection, annotation, and training. The emergence of Vision-Language Models (VLMs) offers a promising alternative, enabling generalization across visual recognition tasks through zero-shot prompting without requiring fine-tuning. This study investigates the potential of VLMs to analyze students' academic emotions via facial expressions in an online learning environment. We employed two VLMs, Llama-3.2-11B-Vision-Instruct and Qwen2.5-VL-7B-Instruct, to analyze 5,000 images depicting confused, distracted, happy, neutral, and tired expressions using zero-shot prompting. Preliminary results indicate that both models demonstrate moderate performance in academic facial expression recognition, with Qwen2.5-VL-7B-Instruct outperforming Llama-3.2-11B-Vision-Instruct. Notably, both models excel in identifying students' happy emotions but fail to detect distracted behavior. Additionally, Qwen2.5-VL-7B-Instruct exhibits relatively high performance in recognizing students' confused expressions, highlighting its potential for practical applications in identifying content that causes student confusion.

情绪识别视觉语言模型在线教育零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。