arXiv:2501.08769cs.CL2025-01被引 19

用合成对话训练心理筛查大模型,准确率超GPT-4

Enhanced Large Language Models for Effective Screening of Depression and Anxiety

  • 构建1157条合成临床对话数据集,支持情绪障碍筛查
  • 在筛查任务中F1达0.7467,解释质量BertScore达0.9408
  • 适合心理健康领域研究者与智能诊疗系统开发者

抑郁症和焦虑症广泛存在,亟需及时识别与管理。近年来大型语言模型(LLMs)提供了潜在解决方案,但高昂的训练成本与数据伦理问题仍存挑战。本文提出一种合成临床访谈的流程,生成1,157条交互式对话(PsyInterview),并推出基于LLM的情绪障碍筛查系统EmoScan。EmoScan可区分粗粒度(如焦虑或抑郁障碍)与细粒度(如重度抑郁障碍)疾病,并开展高质量访谈。评估显示,EmoScan在情绪障碍筛查中超越基础模型及其他LLM(如GPT-4),F1-score达0.7467;其解释质量优异(BERTScore=0.9408),且在外部数据集上表现稳健(F1-score=0.67)。此外,通过自动化评分与人工评估验证,EmoScan在访谈能力上优于基线模型。本工作强调了可扩展数据生成管道对开发有效心理健康大模型工具的重要性。

原文摘要 · Abstract (English)

Depressive and anxiety disorders are widespread, necessitating timely identification and management. Recent advances in Large Language Models (LLMs) offer potential solutions, yet high costs and ethical concerns about training data remain challenges. This paper introduces a pipeline for synthesizing clinical interviews, resulting in 1,157 interactive dialogues (PsyInterview), and presents EmoScan, an LLM-based emotional disorder screening system. EmoScan distinguishes between coarse (e.g., anxiety or depressive disorders) and fine disorders (e.g., major depressive disorders) and conducts high-quality interviews. Evaluations showed that EmoScan exceeded the performance of base models and other LLMs like GPT-4 in screening emotional disorders (F1-score=0.7467). It also delivers superior explanations (BERTScore=0.9408) and demonstrates robust generalizability (F1-score of 0.67 on an external dataset). Furthermore, EmoScan outperforms baselines in interviewing skills, as validated by automated ratings and human evaluations. This work highlights the importance of scalable data-generative pipelines for developing effective mental health LLM tools.

心理筛查大模型对话生成医疗AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。