将视觉信息融入语音语言模型,提升抑郁症检测精度
It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
- 在语音语言模型中加入视觉理解,实现音视频时序对齐
- 在DAIC-WoZ数据集上超越单模态与已有多模态方法
- 可扩展至生理信号,适合临床心理评估场景
抑郁症是全球最普遍的心理健康障碍之一。近年来,语音、视频和文本等多模态数据被广泛用于开发AI辅助抑郁症评估系统。大语言模型凭借强大的语言理解和泛化能力推动了该领域发展。然而,传统大语言模型仍以文本为中心,无法处理音频和视觉模态中的丰富非语言线索,而这些线索在心理健康评估中至关重要。尽管多模态大语言模型前景广阔,但针对心理应用的模型仍寥寥无几。本研究提出一种新型多模态大语言模型框架,用于抑郁症检测。方法通过在语音语言模型中引入视觉理解,并在时间戳级别对齐音视频特征,从而更精细地建模跨模态的时间动态,同时减少对大量训练数据和计算资源的需求。在DAIC-WoZ数据集上的实验表明,该模型优于单模态方法及现有多模态方法。此外,该框架可进一步集成额外生理信号,为心理健康的更广泛应用铺平道路。
原文摘要 · Abstract (English)
Depression is one of the most prevalent mental health disorders globally. In recent years, multi-modal data, such as speech, video, and transcripts, has been increasingly used to develop AI-assisted depression assessment systems. Large language models have further advanced this field due to their strong language understanding and generalization capabilities. However, conventional LLMs remain text-centric and cannot process the rich non-verbal cues found in audio and visual modalities, which are critical components in mental health evaluation. While multi-modal LLMs offer a promising direction, few are tailored for psychological applications. In this study, we propose a novel multi-modal LLM framework for depression detection. Our approach augments an audio language model with visual understanding and aligns audio-visual features at the timestamp level. This fine-grained alignment improves modeling of temporal dynamics across modalities while reducing the need for extensive training data and computational resources. Experiments on the DAIC-WoZ dataset demonstrate that our model outperforms both single-modality approaches and previous multi-modal methods. Moreover, the proposed framework can be extended to incorporate additional physiological signals, paving the way for broader clinical applications beyond mental health.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。