arXiv:2503.06620eess.AS2025-03被引 6

预训练模型在抑郁检测中表现差,因语音与语义特征混杂。

Why Pre-trained Models Fail: Feature Entanglement in Multi-modal Depression Detection

  • 分离语音与语义的高层特征,缓解特征混杂问题。
  • 改进后,语音自监督模型与大语言模型性能显著提升。
  • 适合关注多模态心理健康检测的研究者。

抑郁症是全球重大心理卫生挑战,推动了人工智能辅助检测的研究。尽管语音自监督模型(SSL Models)已被应用于抑郁检测,但在缺乏大量数据增强时表现不佳。大型语言模型(LLMs)虽在多个领域成功,却尚未在多模态抑郁检测中被探索。本文首次构建基于LLM的系统,揭示其处理多模态信息的根本局限。通过系统分析发现,预训练模型性能差的主要原因是:内容与语音的高层特征在模型表示中混杂,难以建立有效决策边界。为此,我们提出一种信息分离框架,解耦这些特征,显著提升了SSL模型与LLM在抑郁检测中的表现。实验验证了该发现,并表明分离特征融合能大幅超越现有方法,为构建更有效的多模态抑郁检测系统提供了新思路。

原文摘要 · Abstract (English)

Depression remains a pressing global mental health issue, driving considerable research into AI-driven detection approaches. While pre-trained models, particularly speech self-supervised models (SSL Models), have been applied to depression detection, they show unexpectedly poor performance without extensive data augmentation. Large Language Models (LLMs), despite their success across various domains, have not been explored in multi-modal depression detection. In this paper, we first establish an LLM-based system to investigate its potential in this task, uncovering fundamental limitations in handling multi-modal information. Through systematic analysis, we discover that the poor performance of pre-trained models stems from the conflation of high-level information, where \textcolor{black}{one of the important reasons is that} high-level features derived from both content and speech are mixed within pre-trained models model representations, making it challenging to establish effective decision boundaries. To address this, we propose an information separation framework that disentangles these features, significantly improving the performance of both SSL models and LLMs in depression detection. Our experiments validate this finding and demonstrate that the integration of separated features yields substantial improvements over existing approaches, providing new insights for developing more effective multi-modal depression detection systems.

抑郁检测多模态特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。