揭穿抑郁症检测基准的评估陷阱,发现模型表现虚高。
A Multi-Probe Audit of Clinical-Interview Depression Detection Benchmarks
- 用严格交叉验证重评E-DAIC数据集,得宏F1=0.723
- 官方排行榜与内部验证结果严重不符,排名重叠为零
- 文本模型在症状密集片段表现突增,音频模型无变化
本文通过四种互补探测方法,对临床访谈抑郁症检测基准进行审计,涵盖DAIC/E-DAIC、CMDC、ANDROIDS、MODMA和PDCH数据集。首先,在严格的留一被试交叉验证下重新评估E-DAIC,轻量级文本+大模型分数融合模型达到宏F1=0.723,为目前已知该协议下的最高值,提供不依赖官方保留集的保守参考。其次,通过96种模型配置测试官方划分是否支持细粒度排行榜,发现开发集交叉验证与官方测试排名仅中度一致:最佳验证配置在官方榜排名第20,官方冠军在验证中排第41,前三名无重叠,且“胜者”仅在32.3%的被试自助样本中排名第一。第三,外部验证了公开的强基线模型在CMDC和ANDROIDS上的近天花板性能,但零样本迁移至外部语料时显著下降。最后,使用基于SRDS标注的对称密集与稀疏访谈片段,压力测试文本与音频模型,结果显示文本得分在症状密集段急剧上升,音频得分几乎不变,五种子实验中文本减音频差距始终为正。
原文摘要 · Abstract (English)
This paper audits benchmark evaluation in clinical-interview depression detection through four complementary probes across DAIC/E-DAIC, CMDC, ANDROIDS, MODMA, and PDCH. First, we re-evaluate E-DAIC under strict subject-disjoint leave-one-subject-out cross-validation. A lightweight hybrid text-plus-LLM-score model reaches macro-F1 = 0.723 - the highest reported under this protocol, to our knowledge - providing a conservative out-of-fold reference point that does not depend on the privileged official holdout. Second, we test whether the E-DAIC official split supports fine-grained leaderboard rankings by sweeping 96 model configurations across modality bundles, pooling strategies, and learners. Development-side cross-validation and official-test rankings align only moderately: the best cross-validation configuration ranks twentieth on the official test, the official-test winner ranks forty-first by cross-validation, top-3 overlap is zero, and the apparent winner is rank-1 in only 32.3% of subject bootstraps. Third, we externally validate strong public CMDC and ANDROIDS baselines that achieve near-ceiling in-domain performance. Zero-shot transfer to external corpora is substantially weaker. Finally, we stress-test E-DAIC text and audio models using paired symptom-dense versus symptom-light interview slices defined by an SRDS-based annotator. Text scores rise sharply on symptom-dense slices, whereas audio scores remain nearly flat; the text-minus-audio gap is positive across all five seeds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。