用强化学习提升抑郁症诊断大模型的推理能力与可解释性
MDD-Thinker: Towards Large Reasoning Models for Major Depressive Disorder Diagnosis
- 结合监督微调与强化学习,增强模型诊断推理能力
- 在真实临床数据上达82.7%准确率,优于主流模型30%以上
- 兼顾准确性与效率,适合临床辅助诊断场景
重度抑郁症(MDD)是全球致残主因之一,现有诊断多依赖主观评估,难以融合多模态临床信息。本文提出MDD-Thinker,一种基于大语言模型的诊断框架,结合监督微调(SFT)与强化学习(RL),提升推理能力与可解释性。基于英国生物银行数据集生成4万条推理样本,并补充1万条公开心理健康数据,构建训练语料。模型在机器学习、深度学习及前沿大模型基线上进行评估。结果表明,MDD-Thinker达到0.8268的准确率与0.8081的F1分数,相较传统方法(如SVM、MLP)和通用大模型,准确率提升29.0%,F1提升38.1%,AUC提升34.8%。模型在推理表现上接近更大规模模型,同时保持计算高效。本研究首次基于大规模真实临床数据构建可解释的推理型大模型,为精神健康智能诊断提供可行路径。
原文摘要 · Abstract (English)
Background Major depressive disorder (MDD) is a leading cause of global disability, yet current diagnostic approaches often rely on subjective assessments and lack the ability to integrate multimodal clinical information. Large language models (LLMs) hold promise for enhancing diagnostic accuracy through advanced reasoning but face challenges in interpretability, hallucination, and reliance on synthetic data. Methods We developed MDD-Thinker, an LLM-based diagnostic framework that integrates supervised fine-tuning (SFT) with reinforcement learning (RL) to strengthen reasoning ability and interpretability. Using the UK Biobank dataset, we generated 40,000 reasoning samples, supplemented with 10,000 samples from publicly available mental health datasets. The model was fine-tuned on these reasoning corpora, and its diagnostic and reasoning performance was evaluated against machine learning, deep learning, and state-of-the-art LLM baselines. Findings MDD-Thinker achieved an accuracy of 0.8268 and F1-score of 0.8081, significantly outperforming traditional baselines such as SVM and MLP, as well as general-purpose LLMs. Incorporating both SFT and RL yielded the greatest improvements, with relative gains of 29.0% in accuracy, 38.1% in F1-score, and 34.8% in AUC. Moreover, the model demonstrated comparable reasoning performance compared to much larger LLMs, while maintaining computational efficiency. Interpretation This study presents the first reasoning-enhanced LLM framework for MDD diagnosis trained on large-scale real-world clinical data. By integrating SFT and RL, MDD-Thinker balances accuracy, interpretability, and efficiency, offering a scalable approach for intelligent psychiatric diagnostics. These findings suggest that reasoning-oriented LLMs can provide clinically reliable support for MDD detection and may inform broader applications in mental health care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。