用大模型生成可解释的抑郁检测报告,分步提升诊断准确率。
Dynamic Summary Generation for Interpretable Multimodal Depression Detection
- 分三阶段逐步分析:筛查、严重程度分类、连续评分,每阶段生成临床摘要。
- 在E-DAIC和CMDC数据集上,准确率和可解释性均优于现有方法。
- 输出简洁人读报告,适合临床辅助诊断与心理评估场景。
抑郁症因社会污名化和主观症状评价难以被有效筛查,导致广泛漏诊。为此,我们提出一种从粗到细的多阶段框架,利用大语言模型(LLMs)实现精准且可解释的抑郁检测。该流程包含二分类筛查、五分类严重程度判断及连续回归预测。每个阶段均由LLM生成更丰富的临床摘要,指导融合文本、音频、视频特征的多模态融合模块,输出带有透明推理依据的预测结果。系统最终将所有摘要整合为一份简明易懂的评估报告。在E-DAIC与CMDC数据集上的实验表明,该方法在准确率与可解释性方面显著优于当前最优基准。
原文摘要 · Abstract (English)
Depression remains widely underdiagnosed and undertreated because stigma and subjective symptom ratings hinder reliable screening. To address this challenge, we propose a coarse-to-fine, multi-stage framework that leverages large language models (LLMs) for accurate and interpretable detection. The pipeline performs binary screening, five-class severity classification, and continuous regression. At each stage, an LLM produces progressively richer clinical summaries that guide a multimodal fusion module integrating text, audio, and video features, yielding predictions with transparent rationale. The system then consolidates all summaries into a concise, human-readable assessment report. Experiments on the E-DAIC and CMDC datasets show significant improvements over state-of-the-art baselines in both accuracy and interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。