arXiv:2410.15500cs.AIcs.SD2024-10中稿 · Interspeech 2024被引 12

用声音转换技术保护老人和病患语音隐私,同时保留关键语调特征。

Anonymising Elderly and Pathological Speech: Voice Conversion Using DDSP and Query-by-Example

  • 结合可微信号处理与查询示例,分离语言、语调和病理特征
  • 在多类语音数据上显著提升可懂度与语调保真度
  • 适合医疗语音匿名化场景,尤其关注老年与病态语音

语音匿名化旨在通过改变语音中的个人标识信息来保护说话人身份,同时保留语言内容。现有方法难以保留老年人和病理语音中特有的语调与语音模式,而这对于远程健康监测至关重要。为弥补这一空白,我们提出一种基于声音转换的方法(DDSP-QbE),结合可微数字信号处理与查询示例技术,并采用新型损失函数训练,实现语言、语调与语音域表征的解耦,使模型能适应罕见语音模式。客观与主观评估表明,该方法在多种数据集、病理类型及说话人条件下,显著优于当前最优的声音转换方法,在可懂度、语调与领域保留方面表现更优,同时保持高质量与说话人匿名性。专家通过分析十二项临床相关领域属性验证了领域保留效果。

原文摘要 · Abstract (English)

Speech anonymisation aims to protect speaker identity by changing personal identifiers in speech while retaining linguistic content. Current methods fail to retain prosody and unique speech patterns found in elderly and pathological speech domains, which is essential for remote health monitoring. To address this gap, we propose a voice conversion-based method (DDSP-QbE) using differentiable digital signal processing and query-by-example. The proposed method, trained with novel losses, aids in disentangling linguistic, prosodic, and domain representations, enabling the model to adapt to uncommon speech patterns. Objective and subjective evaluations show that DDSP-QbE significantly outperforms the voice conversion state-of-the-art concerning intelligibility, prosody, and domain preservation across diverse datasets, pathologies, and speakers while maintaining quality and speaker anonymity. Experts validate domain preservation by analysing twelve clinically pertinent domain attributes.

语音匿名声音转换医疗语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。