arXiv:2508.08591cs.CLcs.AI2025-08被引 2

基于真实叙事的可解释抑郁检测模型,准确率超90%。

DepressLLM: Interpretable domain-adapted language model for depression detection from real-world narratives

  • 用3699篇自述故事训练可解释模型,融合置信度评估机制
  • 在高置信样本上AUC达0.904,显著优于整体表现
  • 适合临床辅助诊断,帮助发现模型与数据局限

大型语言模型的发展推动了诸多应用,但抑郁预测受限于大规模、高质量、严格标注的数据集匮乏。本研究提出DepressLLM,基于包含3,699篇反映幸福与痛苦的自传体叙事构建的新语料库进行训练与评估。该模型提供可解释的抑郁判断,并通过其得分引导的词概率求和(SToPS)模块,在实现更优分类性能的同时给出可靠置信度估计,整体AUC为0.789,而在置信度≥0.95的样本上提升至0.904。为验证鲁棒性,我们在内部数据集(包括每日压力与情绪记录的生态瞬时评估EMA语料库)及公开临床访谈数据上进行了测试。精神科医生对高置信误判案例的审查揭示了模型与数据的关键局限,指明未来改进方向。结果表明,可解释人工智能有助于早期抑郁症诊断,彰显了医疗AI在精神科的应用前景。

原文摘要 · Abstract (English)

Advances in large language models (LLMs) have enabled a wide range of applications. However, depression prediction is hindered by the lack of large-scale, high-quality, and rigorously annotated datasets. This study introduces DepressLLM, trained and evaluated on a novel corpus of 3,699 autobiographical narratives reflecting both happiness and distress. DepressLLM provides interpretable depression predictions and, via its Score-guided Token Probability Summation (SToPS) module, delivers both improved classification performance and reliable confidence estimates, achieving an AUC of 0.789, which rises to 0.904 on samples with confidence $\geq$ 0.95. To validate its robustness to heterogeneous data, we evaluated DepressLLM on in-house datasets, including an Ecological Momentary Assessment (EMA) corpus of daily stress and mood recordings, and on public clinical interview data. Finally, a psychiatric review of high-confidence misclassifications highlighted key model and data limitations that suggest directions for future refinements. These findings demonstrate that interpretable AI can enable earlier diagnosis of depression and underscore the promise of medical AI in psychiatry.

抑郁检测可解释AI语言模型医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。