arXiv:2501.16813cs.CLcs.SD2025-01被引 2

用文本和音频融合提升抑郁症检测准确率

Multimodal Magic Elevating Depression Detection with a Fusion of Text and Audio Intelligence

  • 采用师生架构与多头注意力动态融合文本和音频特征
  • 在DAIC-WOZ数据集上达到99.1%的F1分数
  • 适合关注心理健康与多模态融合研究者

本研究提出一种基于师生架构的创新多模态融合模型,以提升抑郁症分类的准确性。针对传统方法在特征融合与模态权重分配上的局限,引入多头注意力机制与加权多模态迁移学习。基于DAIC-WOZ数据集,由文本和音频教师模型指导的学生融合模型显著提升了分类性能。消融实验表明,该模型在测试集上取得99.1%的F1分数,明显优于单模态及传统方法。该方法有效捕捉了文本与音频特征间的互补性,并动态调整教师模型贡献,增强了泛化能力。实验结果验证了所提框架在处理复杂多模态数据时的鲁棒性与适应性。本研究为抑郁症分析中的多模态大模型学习提供了新范式,深化了对现有模态融合与特征提取方法局限性的理解。

原文摘要 · Abstract (English)

This study proposes an innovative multimodal fusion model based on a teacher-student architecture to enhance the accuracy of depression classification. Our designed model addresses the limitations of traditional methods in feature fusion and modality weight allocation by introducing multi-head attention mechanisms and weighted multimodal transfer learning. Leveraging the DAIC-WOZ dataset, the student fusion model, guided by textual and auditory teacher models, achieves significant improvements in classification accuracy. Ablation experiments demonstrate that the proposed model attains an F1 score of 99. 1% on the test set, significantly outperforming unimodal and conventional approaches. Our method effectively captures the complementarity between textual and audio features while dynamically adjusting the contributions of the teacher models to enhance generalization capabilities. The experimental results highlight the robustness and adaptability of the proposed framework in handling complex multimodal data. This research provides a novel technical framework for multimodal large model learning in depression analysis, offering new insights into addressing the limitations of existing methods in modality fusion and feature extraction.

抑郁症检测多模态融合师生模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。