arXiv:2604.14354eess.AS2026-04被引 2

现有语音抑郁检测模型依赖说话人身份,而非真实抑郁特征。

Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection

论文配图:Who is Speaking or Who is Depressed? A Controlled Study of Speaker Leakage in Speech-Based Depression Detection
图 1 · 摘自论文原文
  • 通过控制训练测试集说话人重叠,验证模型是否真正学习抑郁特征。
  • 无重叠时准确率骤降,说明模型依赖说话人身份而非抑郁信号。
  • 提醒研究者需采用严格去身份评估,避免高估模型临床实用性。

本研究探究语音抑郁检测模型是学习抑郁相关的声学生物标志物,还是依赖说话人身份线索。基于DAIC-WOZ数据集,提出一种控制训练与测试集说话人重叠的数据划分策略,同时保持训练规模不变,并评估三种不同复杂度的模型。结果表明,说话人重叠显著提升性能,而在未见说话人上准确率急剧下降。即使使用领域对抗神经网络,性能差距依然显著。这表明当前语音模型提取的抑郁特征与说话人身份高度纠缠。传统评估协议可能过度估计模型泛化能力与临床价值,因此亟需严格的说话人无关评估。

原文摘要 · Abstract (English)

This study investigates whether speech-based depression detection models learn depression-related acoustic biomarkers or instead rely on speaker identity cues. Using the DAIC-WOZ dataset, we propose a data-splitting strategy that controls speaker overlap between training and test sets while keeping the training size constant, and evaluate three models of varying complexity. Results show that speaker overlap significantly boosts performance, whereas accuracy drops sharply on unseen speakers. Even with a Domain-Adversarial Neural Network, a substantial performance gap remains. These findings indicate that depression-related features extracted by current speech models are highly entangled with speaker identity. Conventional evaluation protocols may therefore overestimate generalization and clinical utility, highlighting the need for strictly speaker-independent evaluation.

抑郁检测语音分析模型泛化数据隔离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。