arXiv:2601.00303cs.CLcs.AI2026-01被引 1

通过解耦语音与语义,提升抑郁检测在伪装情境下的鲁棒性。

DepFlow: Disentangled Speech Generation to Mitigate Semantic Bias in Depression Detection

  • 分三阶段生成解耦语音,分离抑郁特征与语言情感
  • 构建伪装抑郁数据集,使模型在语义与语音不匹配时仍准确识别
  • 适合需要高鲁棒性的临床辅助诊断系统和模拟测试

语音是早期心理健康筛查中可扩展且无侵入性的生物标志物。然而,如DAIC-WOZ等常用抑郁数据集存在语言情感与诊断标签强耦合的问题,导致模型学习语义捷径,影响真实场景下(如伪装抑郁)的鲁棒性。为此,本文提出DepFlow,一种三阶段抑郁条件的文本到语音框架:首先,通过对抗训练的抑郁声学编码器学习与说话人和内容无关的抑郁嵌入,实现有效解耦并保留判别能力(ROC-AUC: 0.693);其次,采用带FiLM调制的流匹配TTS模型注入这些嵌入,实现对抑郁严重程度的可控合成,同时保持内容与说话人身份不变;第三,基于原型的严重程度映射机制实现抑郁连续体上的平滑可解释操控。利用DepFlow,我们构建了面向伪装抑郁的数据增强集(CDoA),将抑郁声学特征与正向/中性语义配对,创造自然数据中罕见的声语不一致样本。在三种抑郁检测架构上评估,CDoA分别提升宏平均F1 9%、12%、5%,显著优于传统增强策略。除提升模型鲁棒性外,DepFlow还为对话系统和受限于伦理与覆盖范围的真实临床数据缺乏的仿真评估提供可控合成平台。

原文摘要 · Abstract (English)

Speech is a scalable and non-invasive biomarker for early mental health screening. However, widely used depression datasets like DAIC-WOZ exhibit strong coupling between linguistic sentiment and diagnostic labels, encouraging models to learn semantic shortcuts. As a result, model robustness may be compromised in real-world scenarios, such as Camouflaged Depression, where individuals maintain socially positive or neutral language despite underlying depressive states. To mitigate this semantic bias, we propose DepFlow, a three-stage depression-conditioned text-to-speech framework. First, a Depression Acoustic Encoder learns speaker- and content-invariant depression embeddings through adversarial training, achieving effective disentanglement while preserving depression discriminability (ROC-AUC: 0.693). Second, a flow-matching TTS model with FiLM modulation injects these embeddings into synthesis, enabling control over depressive severity while preserving content and speaker identity. Third, a prototype-based severity mapping mechanism provides smooth and interpretable manipulation across the depression continuum. Using DepFlow, we construct a Camouflage Depression-oriented Augmentation (CDoA) dataset that pairs depressed acoustic patterns with positive/neutral content from a sentiment-stratified text bank, creating acoustic-semantic mismatches underrepresented in natural data. Evaluated across three depression detection architectures, CDoA improves macro-F1 by 9%, 12%, and 5%, respectively, consistently outperforming conventional augmentation strategies in depression Detection. Beyond enhancing robustness, DepFlow provides a controllable synthesis platform for conversational systems and simulation-based evaluation, where real clinical data remains limited by ethical and coverage constraints.

抑郁检测语音合成数据增强解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。