通过解耦编码与多阶段蒸馏,实现语音隐私保护与情感保留的平衡。
NPU-NTU System for Voice Privacy 2024 Challenge
- 采用解耦神经编解码架构,分步分离说话人身份与语言内容。
- 在2024语音隐私挑战中,隐私保护与情感保留性能最优。
- 适合关注语音隐私、个性化语音合成的研究者与开发者。
说话人匿名化是一种有效的隐私保护技术,可在保留原始语音的语言内容和副语言信息的同时隐藏说话人身份。为建立公平基准并促进说话人匿名化系统比较,语音隐私挑战(VoicePrivacy Challenge, VPC)于2020年和2022年举办,2024年将推出新版。本文介绍我们在VPC 2024中提出的说话人匿名化系统。该系统采用解耦神经编解码架构和串行解耦策略,逐步分离全局说话人身份与时变的语言内容及副语言信息。我们引入多种蒸馏方法以分离语言内容、说话人身份和情绪,包括语义蒸馏、监督说话人蒸馏和帧级情绪蒸馏。基于这些蒸馏结果,通过加权求和一组候选说话人身份与随机生成说话人身份,对原始说话人身份进行匿名化处理。本系统在VPC 2024中实现了隐私保护与情绪保留的最佳权衡。
原文摘要 · Abstract (English)
Speaker anonymization is an effective privacy protection solution that conceals the speaker's identity while preserving the linguistic content and paralinguistic information of the original speech. To establish a fair benchmark and facilitate comparison of speaker anonymization systems, the VoicePrivacy Challenge (VPC) was held in 2020 and 2022, with a new edition planned for 2024. In this paper, we describe our proposed speaker anonymization system for VPC 2024. Our system employs a disentangled neural codec architecture and a serial disentanglement strategy to gradually disentangle the global speaker identity and time-variant linguistic content and paralinguistic information. We introduce multiple distillation methods to disentangle linguistic content, speaker identity, and emotion. These methods include semantic distillation, supervised speaker distillation, and frame-level emotion distillation. Based on these distillations, we anonymize the original speaker identity using a weighted sum of a set of candidate speaker identities and a randomly generated speaker identity. Our system achieves the best trade-off of privacy protection and emotion preservation in VPC 2024.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。