用多层级方法保护痴呆语音检测中的隐私,不牺牲诊断效果。
Multi-Level Privacy-Preserving Dementia Detection from Speech via Targeted Adversarial Obfuscation and Representation Learning

- 在信号层用关键词扰动最大化识别错误,保留关键语音特征。
- 在特征层用反向梯度加噪声,抑制说话人特征但保留疾病信号。
- 隐私保护强且诊断准确,适合医疗语音隐私研究者使用。
用于痴呆检测的语音记录会暴露说话人身份,引发严重隐私问题。现有方法通常只应对单一威胁,无法平衡隐私与实用性。本文提出多层级框架,有效抵御两种窃听风险:在信号层,累积信号攻击(CSA)将扰动集中在关键词对齐区域,使语音识别错误率(WER)达1.00,同时保留关键的语调生物标志;在特征层,采用带互信息引导噪声注入的梯度反转层(GRL),抑制说话人区分性维度,同时保持痴呆相关诊断结构。在DementiaBank Pitt语料库上评估,说话人识别接近随机水平(等错误率EER=0.59,F1=0.003),而痴呆分类性能良好(F1=0.78,AUC=0.86)。
原文摘要 · Abstract (English)
Speech recordings used for dementia detection inherently expose speaker identity, raising critical privacy concerns. Existing methods typically address only singular threats and fail to resolve the privacy--utility trade-off. We propose a multi-level framework designed to neutralize two distinct eavesdropping vectors. At the signal level, a Cumulative Signal Attack (CSA) concentrates perturbations in keyword-aligned regions to maximize transcription error (Word Error Rate WER = 1.00) while preserving vital prosodic biomarkers. At the feature level, a Gradient Reversal Layer (GRL) with Mutual Information (MI)-guided noise injection suppresses speaker-discriminative dimensions while retaining dementia-relevant diagnostic structure. Evaluated on the DementiaBank Pitt Corpus, our framework achieves near-chance speaker identification (Equal Error Rate EER = 0.59, F1 = 0.003) while maintaining strong dementia classification performance (F1 = 0.78, AUC = 0.86).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。