arXiv:2511.15675cs.CVcs.AI2025-11被引 2

用眼动+表情+语音数据,通过多频图网络提升抑郁检测准确率。

MF-GCN: A Multi-Frequency Graph Convolutional Network for Tri-Modal Depression Detection Using Eye-Tracking, Facial, and Acoustic Features

  • 设计多频滤波模块,同时捕捉低频与高频特征信号。
  • 二分类敏感度达0.96,三分类敏感度0.79、特异度0.87。
  • 在中英文数据集上均表现优异,适合临床辅助诊断场景。

抑郁症是全球普遍的精神健康问题,表现为持续低落情绪和快感缺失,但因依赖主观评估而常被漏诊。为此,我们构建了一个包含103名临床确诊者的金标准数据集,融合眼动、音频与视频数据,全面表征抑郁症状。眼动数据量化对负性刺激的注意偏向,音频与视频数据捕捉情感平淡与运动迟缓特征。统计验证显示其具有显著区分能力。针对现有图模型仅关注低频信息的局限,提出多频图卷积网络(MF-GCN),核心为多频滤波器组模块(MFFBM),可同时利用高低频信号。在多种基线模型对比中,MF-GCN持续领先:二分类任务敏感度0.96,F2分数0.94;三分类任务敏感度0.79,特异度0.87,显著优于其他模型。在中文多模态抑郁语料库(CMDC)上测试,敏感度0.95,F2分数0.96。结果表明,该三模态多频框架能有效捕捉跨模态交互,实现精准抑郁检测。

原文摘要 · Abstract (English)

Depression is a prevalent global mental health disorder, characterised by persistent low mood and anhedonia. However, it remains underdiagnosed because current diagnostic methods depend heavily on subjective clinical assessments. To enable objective detection, we introduce a gold standard dataset of 103 clinically assessed participants collected through a tripartite data approach which uniquely integrated eye tracking data with audio and video to give a comprehensive representation of depressive symptoms. Eye tracking data quantifies the attentional bias towards negative stimuli that is frequently observed in depressed groups. Audio and video data capture the affective flattening and psychomotor retardation characteristic of depression. Statistical validation confirmed their significant discriminative power in distinguishing depressed from non depressed groups. We address a critical limitation of existing graph-based models that focus on low-frequency information and propose a Multi-Frequency Graph Convolutional Network (MF-GCN). This framework consists of a novel Multi-Frequency Filter Bank Module (MFFBM), which can leverage both low and high frequency signals. Extensive evaluation against traditional machine learning algorithms and deep learning frameworks demonstrates that MF-GCN consistently outperforms baselines. In binary classification, the model achieved a sensitivity of 0.96 and F2 score of 0.94. For the 3 class classification task, the proposed method achieved a sensitivity of 0.79 and specificity of 0.87 and siginificantly suprassed other models. To validate generalizability, the model was also evaluated on the Chinese Multimodal Depression Corpus (CMDC) dataset and achieved a sensitivity of 0.95 and F2 score of 0.96. These results confirm that our trimodal, multi frequency framework effectively captures cross modal interaction for accurate depression detection.

抑郁检测多模态图神经网络眼动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。