arXiv:2607.02920eess.AS2026-07

跨语言抑郁检测新方法,避免身份泄露提升真实性能

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment

论文配图:Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment
图 1 · 摘自论文原文
  • 用对比对齐将英汉语音嵌入映射到共享临床空间
  • 在52名说汉语者上实现F1 0.640,优于基线0.622
  • 揭示此前高分源于身份泄露,适合多语言心理健康研究

不同语言群体的抑郁症诊断与临床表现存在显著差异。基于语音的抑郁检测在单语言场景下表现良好,但跨语言泛化仍是难题。主要原因在于以往研究采用段落级随机划分且未按说话人分组,导致身份信息泄露,虚高了评估指标。本文提出CLeaD框架,一种无需平行数据或目标语言微调的监督对比对齐方法,将英语和中文的WavLM嵌入映射至共享临床空间。在52名汉语说话人上,对比对齐在留一说话人外评估中略优于基线(F1: 0.640 vs. 0.622),并在中间层(第7-8层)提升抑郁类召回率。两个发现保持稳健:模型规模扩大虽提升英文单语性能,却恶化跨语言表现;此前报道的汉语F1高达0.954是身份泄露所致的人为结果,我们复现并量化了该偏差。

原文摘要 · Abstract (English)

Significant disparities exist in the diagnosis and clinical presentation of depression across different linguistic populations. Speech-based depression detection performs well monolingually, but cross-lingual generalization remains an open challenge. A key reason is that prior work uses segment-level random splits without speaker grouping, leading to identity leakage that inflates reported metrics. We propose CLeaD, a supervised contrastive alignment framework that maps WavLM embeddings from English and Mandarin into a shared clinical space, without parallel data or target-language fine-tuning. Evaluating 52 Mandarin speakers, contrastive alignment modestly outperforms the baseline (F1: 0.640 vs. 0.622) under leave-one-speaker-out evaluation. It also improves depressed-class recall at intermediate layers (7-8), though the small test set limits generalizability. Two findings remain robust: model scaling degrades cross-lingual performance while improving monolingual English, and speaker identity leakage artificially inflated previously reported Mandarin F1 scores to 0.954, an artifact we reproduce and quantify.

跨语言语音分析抑郁检测对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。