让失智症患者说话更清晰,同时保护隐私。
ClaritySpeech: Dementia Obfuscation in Speech
- 用ASR+文本混淆+零样本语音合成,修复失智症语音
- 在低数据下仍保持50%说话人相似度,误识率大幅下降
- 适合医疗辅助与失智症患者隐私保护场景
失智症会改变说话模式,造成沟通障碍并引发隐私担忧。当前语音技术(如自动语音识别,ASR)难以处理失智症或异常语音,进一步影响可及性。本文提出一种名为ClaritySpeech的新框架,融合ASR、文本混淆和零样本文语转换(TTS),在无需微调的低数据环境下,修复失智症患者的语音并保留说话人身份。实验显示,在多种对抗性设置与模态(音频、文本、融合)下,ADReSS与ADReSSo数据集的平均F1分数分别下降16%和10%,但说话人相似度仍维持在50%。系统还将语音识别错误率显著降低(ADReSS从0.73降至0.08,ADReSSo从0.15降至0.08),语音质量从1.65提升至约2.15,有效增强隐私保护与交流可及性。
原文摘要 · Abstract (English)
Dementia, a neurodegenerative disease, alters speech patterns, creating communication barriers and raising privacy concerns. Current speech technologies, such as automatic speech transcription (ASR), struggle with dementia and atypical speech, further challenging accessibility. This paper presents a novel dementia obfuscation in speech framework, ClaritySpeech, integrating ASR, text obfuscation, and zero-shot text-to-speech (TTS) to correct dementia-affected speech while preserving speaker identity in low-data environments without fine-tuning. Results show a 16% and 10% drop in mean F1 score across various adversarial settings and modalities (audio, text, fusion) for ADReSS and ADReSSo, respectively, maintaining 50% speaker similarity. We also find that our system improves WER (from 0.73 to 0.08 for ADReSS and 0.15 for ADReSSo) and speech quality from 1.65 to ~2.15, enhancing privacy and accessibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。