用门诊对话自动识别抑郁,效果接近人工筛查。
Depression Detection at the Point of Care: Automated Analysis of Linguistic Signals from Routine Primary Care Encounters
- 从医生患者对话中提取语言特征,用AI判断抑郁状态。
- 首次128个患者词就能达到较好检测效果(AUPRC=0.356)。
- 适合临床实时辅助,无需额外负担患者或医生。
抑郁症在初级医疗中常被漏诊,但及时发现至关重要。随着数字记录技术普及,临床就诊录音为自然对话中的抑郁检测提供了可能。本研究基于1,108段来自Establishing Focus研究的录音,以PHQ-9作为抑郁判定标准(确诊253例,非确诊855例),比较了三种监督学习方法(Sentence-BERT+LR、LIWC+LR、ModernBERT)与零样本GPT-OSS的表现。结果表明,GPT-OSS表现最佳(AUPRC=0.510,AUROC=0.774),LIWC+LR在监督模型中也表现良好(AUPRC=0.500,AUROC=0.742)。联合医患双人对话文本优于单一说话人配置,因医生在抑郁患者对话中会模仿其语言模式,形成叠加信号。仅需前128个患者词即可实现有意义检测(AUPRC=0.356,AUROC=0.675),支持即时临床决策支持。研究建议将被动采集的临床音频作为现有筛查流程的低负担补充。
原文摘要 · Abstract (English)
Depression is underdiagnosed in primary care, yet timely identification remains critical. Recorded clinical encounters, increasingly common with digital scribing technologies, present an opportunity to detect depression from naturalistic dialogue. We investigated automated depression detection from 1,108 audio-recorded primary care encounters in the Establishing Focus study, with depression defined by PHQ-9 (n=253 depressed, n=855 non-depressed). We compared three supervised approaches, Sentence-BERT + Logistic Regression (LR), LIWC+LR and ModernBERT, against a zero-shot GPT-OSS. GPT-OSS achieved the strongest performance (AUPRC=0.510, AUROC=0.774), with LIWC+LR competitive among supervised models (AUPRC=0.500, AUROC=0.742). Combined dyadic transcripts outperformed single-speaker configurations, with providers linguistically mirroring patients in depression encounters, an additive signal not captured by either speaker alone. Meaningful detection is achievable from the first 128 patient tokens (AUPRC=0.356, AUROC=0.675), supporting in-the-moment clinical decision support. These findings argue for passively collected clinical audio as a low-burden complement to existing screening workflows.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。