用互注意力融合音视频特征,提升抑郁症检测准确率
MDD-Net: Multimodal Depression Detection through Mutual Transformer
- 设计互注意力机制,同步提取并融合音频与视觉特征
- 在D-Vlog数据集上F1分数比现有方法最高提升17.37%
- 适合心理健康监测、多模态情感分析方向的研究者
抑郁症是严重损害个体情绪与身体健康的常见心理疾病。社交媒体平台的数据采集便捷性激发了其在心理健康研究中的应用兴趣。本文提出一种多模态抑郁症检测网络(MDD-Net),利用来自社交媒体的声学与视觉数据,通过互变压器高效提取并融合多模态特征以实现精准抑郁检测。MDD-Net包含四个核心模块:声学特征提取模块用于获取相关声学属性,视觉特征提取模块用于提取显著高层模式,互变压器用于计算生成特征间的相关性并融合多模态特征,检测层则基于融合特征表示进行抑郁症识别。在多模态D-Vlog数据集上的大量实验表明,所提出的多模态抑郁症检测网络在F1分数上相比现有最优方法最高提升17.37%,验证了系统的优越性能。源代码可在https://github.com/rezwanh001/Multimodal-Depression-Detection 获取。
原文摘要 · Abstract (English)
Depression is a major mental health condition that severely impacts the emotional and physical well-being of individuals. The simple nature of data collection from social media platforms has attracted significant interest in properly utilizing this information for mental health research. A Multimodal Depression Detection Network (MDD-Net), utilizing acoustic and visual data obtained from social media networks, is proposed in this work where mutual transformers are exploited to efficiently extract and fuse multimodal features for efficient depression detection. The MDD-Net consists of four core modules: an acoustic feature extraction module for retrieving relevant acoustic attributes, a visual feature extraction module for extracting significant high-level patterns, a mutual transformer for computing the correlations among the generated features and fusing these features from multiple modalities, and a detection layer for detecting depression using the fused feature representations. The extensive experiments are performed using the multimodal D-Vlog dataset, and the findings reveal that the developed multimodal depression detection network surpasses the state-of-the-art by up to 17.37% for F1-Score, demonstrating the greater performance of the proposed system. The source code is accessible at https://github.com/rezwanh001/Multimodal-Depression-Detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。