融合视觉与生理信号,提升学习者专注度识别准确率。
VisioPhysioENet: Visual Physiological Engagement Detection Network
- 双层特征提取:用Dlib和OpenCV捕捉面部特征,结合皮肤正交法分析生理信号。
- 在DAiSEE数据集上达到63.09%准确率,优于仅用多模态的现有模型8.6%。
- 适合教育技术、人机交互领域研究者关注注意力检测新方法。
本文提出VisioPhysioENet,一种利用视觉与生理信号联合检测学习者参与度的新型多模态系统。采用两级特征提取策略:视觉特征通过Dlib检测人脸关键点,搭配OpenCV进行补充估计;基于Dlib构建的人脸识别库专门定位用于生理信号提取的感兴趣区域。生理信号则通过皮肤正交法评估心血管活动。上述特征经由先进机器学习分类器融合,有效提升对不同参与度水平的识别能力。在DAiSEE数据集上全面测试,该模型取得63.09%的准确率,显著优于多数现有方法。相较于唯一同时使用生理与视觉特征的其他模型,性能提升8.6%。
原文摘要 · Abstract (English)
This paper presents VisioPhysioENet, a novel multimodal system that leverages visual and physiological signals to detect learner engagement. It employs a two-level approach for extracting both visual and physiological features. For visual feature extraction, Dlib is used to detect facial landmarks, while OpenCV provides additional estimations. The face recognition library, built on Dlib, is used to identify the facial region of interest specifically for physiological signal extraction. Physiological signals are then extracted using the plane-orthogonal-toskin method to assess cardiovascular activity. These features are integrated using advanced machine learning classifiers, enhancing the detection of various levels of engagement. We thoroughly tested VisioPhysioENet on the DAiSEE dataset. It achieved an accuracy of 63.09%. This shows it can better identify different levels of engagement compared to many existing methods. It performed 8.6% better than the only other model that uses both physiological and visual features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。