构建高质量困惑识别基准,助力教育视频精准诊断学习障碍
ConfusionBench: An Expert-Validated Benchmark for Confusion Recognition and Localization in Educational Videos
- 三阶段过滤流程:模型初筛+研究者精修+专家验证,提升数据质量
- 包含片段级识别与长视频定位双任务数据集,支持细粒度分析
- 提供可视化报告,帮助教师制定干预策略,适合教育AI研究者使用
从视频中识别并定位学生困惑是教育人工智能的重要挑战。现有困惑数据集存在标签噪声、粗粒度时间标注和专家验证不足等问题,制约了细粒度识别与时间对齐分析的可靠性。为此,我们提出一种多阶段过滤流程,融合两阶段模型辅助筛选、研究者人工校正及专家验证,构建更高质量的困惑理解基准。基于此流程,我们推出ConfusionBench,一个由平衡的困惑识别数据集与视频定位数据集组成的教育视频新基准。我们还对代表性开源模型与专有模型进行了零样本基线评估,结果表明:专有模型整体表现更优但易过度预测过渡段,开源模型则更保守且漏检率更高。此外,提出的困惑报告可视化工具可辅助教育专家决策干预措施与调整学习计划。所有数据集及相关材料将公开发布于项目主页。
原文摘要 · Abstract (English)
Recognizing and localizing student confusion from video is an important yet challenging problem in educational AI. Existing confusion datasets suffer from noisy labels, coarse temporal annotations, and limited expert validation, which hinder reliable fine-grained recognition and temporally grounded analysis. To address these limitations, we propose a practical multi-stage filtering pipeline that integrates two stages of model-assisted screening, researcher curation, and expert validation to build a higher-quality benchmark for confusion understanding. Based on this pipeline, we introduce ConfusionBench, a new benchmark for educational videos consisting of a balanced confusion recognition dataset and a video localization dataset. We further provide zero-shot baseline evaluations of a representative open-source model and a proprietary model on clip-level confusion recognition, long-video confusion localization tasks. Experimental results show that the proprietary model performs better overall but tends to over-predict transitional segments, while the open-source model is more conservative and more prone to missed detections. In addition, the proposed student confusion report visualization can support educational experts in making intervention decisions and adapting learning plans accordingly. All datasets and related materials will be made publicly available on our project page.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。