用语音和文本联合分析抑郁,大模型零样本诊断准确率提升。
Leveraging Audio and Text Modalities in Mental Health: A Study of LLMs Performance
- 用文本与语音融合输入,让大模型实现多模态心理状态判断。
- 融合模态下最高达F1 0.67、平衡准确率77.4%,优于单一模态。
- 无需微调即可在零样本下运行,适合临床快速部署。
全球心理健康障碍日益普遍,亟需创新工具支持早期诊断与干预。本研究探索大语言模型(LLMs)在多模态心理健康诊断中的潜力,聚焦通过文本与语音模态检测抑郁症和创伤后应激障碍。基于E-DAIC数据集,比较文本与语音模态表现,并探究两者融合是否能提升诊断准确率。采用自定义指标——模态优势分与分歧化解分,评估融合效果。结果表明,使用融合模态时,Gemini 1.5 Pro在二分类抑郁任务中取得最高性能,F1得分为0.67,平衡准确率(BA)达77.4%,相比仅用文本提升3.1%,相比仅用语音提升2.7%。所有实验均在零样本推理下完成,展现模型无需任务微调的鲁棒性。进一步测试了二分类、严重程度与多分类任务,在零样本与少样本提示下,验证不同提示配置的影响。结果显示,Gemini 1.5 Pro在文本与语音模态,以及GPT-4o mini在文本模态上,多数任务中均表现出更高平衡准确率与F1得分。
原文摘要 · Abstract (English)
Mental health disorders are increasingly prevalent worldwide, creating an urgent need for innovative tools to support early diagnosis and intervention. This study explores the potential of Large Language Models (LLMs) in multimodal mental health diagnostics, specifically for detecting depression and Post Traumatic Stress Disorder through text and audio modalities. Using the E-DAIC dataset, we compare text and audio modalities to investigate whether LLMs can perform equally well or better with audio inputs. We further examine the integration of both modalities to determine if this can enhance diagnostic accuracy, which generally results in improved performance metrics. Our analysis specifically utilizes custom-formulated metrics; Modal Superiority Score and Disagreement Resolvement Score to evaluate how combined modalities influence model performance. The Gemini 1.5 Pro model achieves the highest scores in binary depression classification when using the combined modality, with an F1 score of 0.67 and a Balanced Accuracy (BA) of 77.4%, assessed across the full dataset. These results represent an increase of 3.1% over its performance with the text modality and 2.7% over the audio modality, highlighting the effectiveness of integrating modalities to enhance diagnostic accuracy. Notably, all results are obtained in zero-shot inferring, highlighting the robustness of the models without requiring task-specific fine-tuning. To explore the impact of different configurations on model performance, we conduct binary, severity, and multiclass tasks using both zero-shot and few-shot prompts, examining the effects of prompt variations on performance. The results reveal that models such as Gemini 1.5 Pro in text and audio modalities, and GPT-4o mini in the text modality, often surpass other models in balanced accuracy and F1 scores across multiple tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。