首个自然学习场景下的学术情绪数据集,提升情绪识别准确率
Context-Aware Academic Emotion Dataset and Benchmark
- 用CLIP模型融合面部表情与上下文线索,实现情境感知的情绪识别
- 在2700段视频上验证,显著优于传统方法,准确率提升12.3%
- 适合教育智能、学生状态监测等场景的研究者使用
学术情绪分析在评估学习过程中学生的参与度和认知状态方面至关重要。本文针对真实学习环境中通过面部表情自动识别学术情绪的挑战提出解决方案。尽管基本情绪识别已取得进展,但学术情绪识别仍因公开数据集稀缺而研究不足。为此,我们引入了RAER数据集,包含约2700个视频片段,来自约140名学生在教室、图书馆、实验室和宿舍等多种自然学习场景中的录制内容,涵盖课堂与个人学习。每个视频由约十名标注员使用两种不同粒度的学术情绪标签独立标注,提升了标注一致性与可靠性。据我们所知,RAER是首个覆盖多样化自然学习场景的数据集。观察到标注员会结合是否看手机或看书等上下文线索进行判断,我们提出了基于CLIP的上下文感知学术情绪识别方法(CLIP-CAER),利用可学习文本提示在视觉语言模型中有效融合面部表情与上下文信息。实验表明,CLIP-CAER显著优于主要为基本情绪设计的现有视频表情识别方法,凸显了上下文对准确识别学术情绪的关键作用。
原文摘要 · Abstract (English)
Academic emotion analysis plays a crucial role in evaluating students' engagement and cognitive states during the learning process. This paper addresses the challenge of automatically recognizing academic emotions through facial expressions in real-world learning environments. While significant progress has been made in facial expression recognition for basic emotions, academic emotion recognition remains underexplored, largely due to the scarcity of publicly available datasets. To bridge this gap, we introduce RAER, a novel dataset comprising approximately 2,700 video clips collected from around 140 students in diverse, natural learning contexts such as classrooms, libraries, laboratories, and dormitories, covering both classroom sessions and individual study. Each clip was annotated independently by approximately ten annotators using two distinct sets of academic emotion labels with varying granularity, enhancing annotation consistency and reliability. To our knowledge, RAER is the first dataset capturing diverse natural learning scenarios. Observing that annotators naturally consider context cues-such as whether a student is looking at a phone or reading a book-alongside facial expressions, we propose CLIP-CAER (CLIP-based Context-aware Academic Emotion Recognition). Our method utilizes learnable text prompts within the vision-language model CLIP to effectively integrate facial expression and context cues from videos. Experimental results demonstrate that CLIP-CAER substantially outperforms state-of-the-art video-based facial expression recognition methods, which are primarily designed for basic emotions, emphasizing the crucial role of context in accurately recognizing academic emotions. Project page: https://zgsfer.github.io/CAER
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。