arXiv:2510.08593cs.CLcs.AI2025-10被引 2

通过分层交互建模语音中的声学与语义特征,提升抑郁症检测准确率

Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech

  • 构建双流表示框架,用异步交叉注意力对齐声学与语义信息
  • 在DAIC-WOZ和MODMA数据集上达到0.81和0.82的宏平均F1分数
  • 无需逐帧标注,利用CTC辅助监督处理抑郁特征的时间不规则分布

基于语音的抑郁症检测(SDD)已成为一种非侵入且可扩展的替代传统临床评估的方法。然而,现有方法仍难以捕捉稀疏且异质的抑郁相关语音特征。尽管预训练自监督学习(SSL)模型提供丰富表征,但多数近期研究仅从单一层级提取特征用于下游分类器,忽略了不同层级中蕴含的低层声学特征与高层语义信息之间的互补性。为显式建模话语内声学与语义表示的交互,我们提出一种带有先验知识的分层自适应表示编码器,通过不对称交叉注意力解耦并重新对齐声学与语义信息,使细粒度声学模式能在语义上下文中被解释。此外,引入连接时序分类(CTC)目标作为辅助监督,以应对抑郁特征在时间上的不规则分布,且无需逐帧标注。在DAIC-WOZ和MODMA数据集上的实验表明,HAREN-CTC在性能上限评估和泛化评估设置下均显著优于现有方法,分别实现0.81和0.82的宏平均F1得分,并在严格交叉验证下保持统计显著的精度与AUC提升。结果表明,建模分层声学-语义交互更真实反映抑郁特征在自然语音中的表现,推动可扩展、客观的抑郁症评估。

原文摘要 · Abstract (English)

Speech-based depression detection (SDD) has emerged as a non-invasive and scalable alternative to conventional clinical assessments. However, existing methods still struggle to capture robust depression-related speech characteristics, which are sparse and heterogeneous. Although pretrained self-supervised learning (SSL) models provide rich representations, most recent SDD studies extract features from a single layer of the pretrained SSL model for the downstream classifier. This practice overlooks the complementary roles of low-level acoustic features and high-level semantic information inherently encoded in different SSL model layers. To explicitly model interactions between acoustic and semantic representations within an utterance, we propose a hierarchical adaptive representation encoder with prior knowledge that disengages and re-aligns acoustic and semantic information through asymmetric cross-attention, enabling fine-grained acoustic patterns to be interpreted in semantic context. In addition, a Connectionist Temporal Classification (CTC) objective is applied as auxiliary supervision to handle the irregular temporal distribution of depressive characteristics without requiring frame-level annotations. Experiments on DAIC-WOZ and MODMA demonstrate that HAREN-CTC consistently outperforms existing methods under both performance upper-bound evaluation and generalization evaluation settings, achieving Macro F1 scores of 0.81 and 0.82 respectively in upper-bound evaluation, and maintaining superior performance with statistically significant improvements in precision and AUC under rigorous cross-validation. These findings suggest that modeling hierarchical acoustic-semantic interactions better reflects how depressive characteristics manifest in natural speech, enabling scalable and objective depression assessment.

抑郁症检测自监督学习语音分析分层建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。