arXiv:2410.02371eess.AScs.AI2024-10被引 13

改进基线模型,在保护语音隐私的同时提升音质与匿名效果。

NTU-NPU System for Voice Privacy 2024 Challenge

  • 用情绪嵌入和多种声纹编码器增强基线模型
  • 提出均值回归基频技术,提升隐私性且不损失语音可用性
  • 探索解耦模型,适合语音隐私保护研究者参考

本文描述我们针对2024年语音隐私挑战赛的参赛方案。未提出全新语音匿名系统,而是对给定基线进行优化以满足所有要求并提升评估指标。具体包括:在B3基线上引入情绪嵌入,并实验WavLM与ECAPA2声纹嵌入器;对比不同声纹与语调匿名化技术;在B5基线上引入均值回归基频(Mean Reversion F0),有效增强隐私性而无性能损失;最后探索$β$-VAE与NaturalSpeech3 FACodec等解耦模型。

原文摘要 · Abstract (English)

In this work, we describe our submissions for the Voice Privacy Challenge 2024. Rather than proposing a novel speech anonymization system, we enhance the provided baselines to meet all required conditions and improve evaluated metrics. Specifically, we implement emotion embedding and experiment with WavLM and ECAPA2 speaker embedders for the B3 baseline. Additionally, we compare different speaker and prosody anonymization techniques. Furthermore, we introduce Mean Reversion F0 for B5, which helps to enhance privacy without a loss in utility. Finally, we explore disentanglement models, namely $β$-VAE and NaturalSpeech3 FACodec.

语音隐私声纹匿名解耦学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。