arXiv:2510.14937cs.CL2025-10

用AI分析临床对话,提前精准识别抑郁焦虑和创伤后应激障碍。

AI-Powered Early Diagnosis of Mental Health Disorders from Real-World Clinical Conversations

  • 用GPT-4.1 Mini、MetaLLaMA和微调RoBERTa模型分析真实对话
  • 对创伤后应激障碍诊断准确率达89%,召回率高达98%
  • 轻量微调(LoRA rank 8/16)即可高效提升检测能力

精神健康障碍是全球致残的主要原因,但抑郁症、焦虑症和创伤后应激障碍(PTSD)常因主观评估、资源有限及污名化而漏诊或误诊。初级医疗中,超过60%的抑郁或焦虑病例被误判,亟需可扩展、易获取且上下文敏感的辅助诊断工具。本研究基于553段真实半结构化访谈数据,每段配以确诊结果,评估多种机器学习模型在主要抑郁发作(MDE)、焦虑障碍和PTSD筛查中的表现。对比包括GPT-4.1 Mini和MetaLLaMA的零样本提示,以及使用低秩适应(LoRA)微调的RoBERTa模型。所有模型在各类诊断中均达80%以上准确率,其中对PTSD的准确率最高达89%,召回率高达98%。研究还发现,聚焦于短片段上下文可提升召回率,表明关键叙事线索有助于提高检测灵敏度。LoRA微调在低秩设置(如秩8和16)下仍保持优异性能,兼具效率与效果。结果表明,基于大模型的方法显著优于传统自评筛查工具,为低门槛、智能早筛提供可能。该工作为将机器学习融入现实临床流程奠定基础,尤其适用于资源匮乏或高污名化环境中。

原文摘要 · Abstract (English)

Mental health disorders remain among the leading cause of disability worldwide, yet conditions such as depression, anxiety, and Post-Traumatic Stress Disorder (PTSD) are frequently underdiagnosed or misdiagnosed due to subjective assessments, limited clinical resources, and stigma and low awareness. In primary care settings, studies show that providers misidentify depression or anxiety in over 60% of cases, highlighting the urgent need for scalable, accessible, and context-aware diagnostic tools that can support early detection and intervention. In this study, we evaluate the effectiveness of machine learning models for mental health screening using a unique dataset of 553 real-world, semistructured interviews, each paried with ground-truth diagnoses for major depressive episodes (MDE), anxiety disorders, and PTSD. We benchmark multiple model classes, including zero-shot prompting with GPT-4.1 Mini and MetaLLaMA, as well as fine-tuned RoBERTa models using LowRank Adaptation (LoRA). Our models achieve over 80% accuracy across diagnostic categories, with especially strongperformance on PTSD (up to 89% accuracy and 98% recall). We also find that using shorter context, focused context segments improves recall, suggesting that focused narrative cues enhance detection sensitivity. LoRA fine-tuning proves both efficient and effective, with lower-rank configurations (e.g., rank 8 and 16) maintaining competitive performance across evaluation metrics. Our results demonstrate that LLM-based models can offer substantial improvements over traditional self-report screening tools, providing a path toward low-barrier, AI-powerd early diagnosis. This work lays the groundwork for integrating machine learning into real-world clinical workflows, particularly in low-resource or high-stigma environments where access to timely mental health care is most limited.

心理健康AI诊断大模型早筛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。