用信息论方法保护语音中的隐私,同时保持抑郁检测准确率
InfoShield: Privacy-Preserving Speech Representations for Mental Health Screening via Information-Theoretic Optimization
- 通过最小化语音表示与敏感属性的互信息实现隐私保护
- 在安卓语料库上将性别识别率从92.6%降至55.5%,年龄识别率从55.7%降至30.3%
- 适合关注语音隐私与心理健康筛查平衡的研究者
基于语音的精神健康筛查可实现抑郁的规模化检测,但临床应用面临用户对个人身份信息泄露的担忧。现有方法难以解决这一矛盾:对抗训练对未知威胁无效,差分隐私则因向所有特征注入噪声而损害诊断性能。本文提出InfoShield,通过信息论优化,在保留抑郁分类准确率的同时最小化语音表示与敏感属性间的互信息。我们发现标准MINE估计器在序列语音中因时间-静态错位而表现不佳,因此引入TimeAwareMINE,采用跨模态注意力对齐声学帧与属性嵌入。在Androids Corpus上的实验表明,InfoShield将性别推断准确率从92.6%降至55.5%,年龄推断从55.7%降至30.3%,仅损失6%的F1(F1=0.784),优于先前最优模型的0.723。
原文摘要 · Abstract (English)
Speech-based mental health screening offers scalable depression detection, yet clinical deployment faces a significant barrier: users' privacy concerns about demographic information exposure. Current techniques struggle to resolve this conflict. Adversarial training often fails against unseen threats, whereas Differential Privacy tends to compromise diagnostic performance by injecting noise across all features. This paper presents InfoShield, which minimizes mutual information between speech representations and sensitive attributes while preserving depression classification accuracy. We identify that standard MINE estimators struggle with sequential speech due to temporal-static misalignment, and introduce TimeAwareMINE with cross-modal attention to align acoustic frames with attribute embeddings. Experiments on the Androids Corpus show InfoShield reduces gender inference from 92.6\% to 55.5\% and age inference from 55.7\% to 30.3\% with limited utility loss (6\% F1 reduction), achieving F1=0.784 compared to prior SOTA's 0.723.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。