让抑郁症识别结果可解释,医生能看懂AI的判断依据。
Explainable Multimodal Depression Recognition in Clinical Interviews via PHQ-Aligned Symptom Summarization
- 用症状摘要融合多模态数据,模拟临床诊断流程。
- 在DAIC-WOZ基础上构建带PHQ-8标注的新数据集。
- 模型输出可读摘要,适合临床医生审查与信任。
近期多模态抑郁症识别研究在临床访谈中整合文本、语音和面部信息,展现了人工智能的潜力。但现有方法对可解释性关注不足,限制了可复现性和医生评审。为此,我们提出 Explain-MDRC 框架,通过生成结构化症状摘要并融合非语言线索实现可解释识别。具体地,我们基于 DAIC-WOZ 构建了 Explain-DAIC 数据集,并加入 PHQ-8 对齐的摘要标注,为可解释模型提供基础。进一步提出 PhqCML 模型,结合 PHQ-8 对齐的症状摘要、感知 PHQ-8 的对比学习以及摘要引导的多模态融合。自动评估与专家评测均表明,Explain-MDRC 提升了识别性能,并生成更具可读性的中间证据,为透明化抑郁症辅助识别提供了新方向。
原文摘要 · Abstract (English)
Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues. However, existing methods pay limited attention to interpretability, thereby constraining reproducibility and clinician review. To address this, we introduce Explain-MDRC, an explainable MDRC framework that mirrors clinical workflows by generating structured symptom summaries from text and integrating them with nonverbal cues for recognition. Specifically, we construct Explain-DAIC, a dataset based on DAIC-WOZ and enriched with PHQ-8-aligned summary annotations, providing a foundation for developing models with built-in interpretability. We further propose PhqCML, a model that combines PHQ-8-aligned symptom summarization with PHQ-aware contrastive learning and summary-informed multimodal fusion. Automated metrics and expert evaluations show that Explain-MDRC improves recognition performance and provides more interpretable, clinician-readable intermediate evidence, suggesting a promising direction for transparent AI-assisted depression recognition research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。