arXiv:2608.26148cs.CLcs.SD2026-08

用语音特征解释抑郁症状,让诊断更透明可懂。

Towards Interpretable Depression Detection: Linking Acoustic Features to DSM-5 Indicators

  • 构建声学特征与临床指标的显式映射框架
  • 在DAIC-WOZ数据集上验证了特征与症状的一致性
  • 本地运行保障隐私,适合医疗场景落地

抑郁症影响全球数百万人,但诊断依赖主观自述,可能遗漏真实行为表现。本文提出一种透明的关联框架,将语音声学特征(如音高变化、停顿、语速)与DSM-5抑郁行为指标进行显式关联,实现可解释的指标级输出。系统可在普通硬件上本地运行,保障用户隐私。在DAIC-WOZ数据集上的初步评估显示,声学特征与心理运动改变及注意力困难等DSM-5指标呈现方向一致的关联,验证了设计合理性。未来工作将基于纵向数据集进一步验证,并在保持边缘计算约束的前提下拓展多模态融合。

原文摘要 · Abstract (English)

Depression affects millions worldwide, yet diagnosis relies on subjective self-reports that may miss authentic behavior. This paper presents an approach linking speech acoustics to DSM-5 depressive-behavior indicators through a transparent Linkage Framework. Unlike black-box models, the framework explicitly maps acoustic features (pitch variability, pauses, speech tempo) to clinical indicators, enabling interpretable, indicator-level outputs. The system runs locally on commodity hardware (HW) to preserve privacy. Preliminary evaluation on DAIC-WOZ shows directionally consistent associations between acoustic features and DSM-5 indicators for psychomotor change and concentration difficulty, supporting the design rationale. Future work will validate on longitudinal datasets and extend multimodal integration while maintaining edge constraints.

抑郁症检测语音分析可解释性隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。