语音匿名化难以同时保护隐私与保留情绪,二者存在本质冲突。
Privacy versus Emotion Preservation Trade-offs in Emotion-Preserving Speaker Anonymization
- 设计多种语音匿名化流程,测试隐私与情绪保留的权衡
- 仅用情绪特征即可训练出半有效的说话人验证系统
- 需专用情绪识别器才能兼顾隐私与情绪保持,适合语音安全研究者
语音技术的进步使个人身份信息极易通过语音泄露。为保护此类信息,差分隐私领域探索了在保留语言和副语言特征的前提下匿名化语音的方法。然而,在保持说话人情绪状态方面仍面临挑战。本文在 VoicePrivacy 2024 挑战背景下研究该问题,开发了多种说话人匿名化流程,发现现有方法要么擅长匿名化,要么擅长保留情绪,难以同时实现。实现双重目标需依赖领域内的情绪识别器。此外,我们发现仅使用情绪表征即可训练出半有效的说话人验证系统,表明情绪与说话人身份特征难以分离,凸显了隐私与情绪保留之间的根本矛盾。
原文摘要 · Abstract (English)
Advances in speech technology now allow unprecedented access to personally identifiable information through speech. To protect such information, the differential privacy field has explored ways to anonymize speech while preserving its utility, including linguistic and paralinguistic aspects. However, anonymizing speech while maintaining emotional state remains challenging. We explore this problem in the context of the VoicePrivacy 2024 challenge. Specifically, we developed various speaker anonymization pipelines and find that approaches either excel at anonymization or preserving emotion state, but not both simultaneously. Achieving both would require an in-domain emotion recognizer. Additionally, we found that it is feasible to train a semi-effective speaker verification system using only emotion representations, demonstrating the challenge of separating these two modalities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。