arXiv:2607.03985eess.AScs.AI2026-07

用变分自编码器生成多样伪说话人,保护语音隐私同时保持音色自然。

NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization

论文配图:NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
图 1 · 摘自论文原文
  • 基于层次化变分自编码器生成高表达力伪说话人嵌入
  • 在语音验证攻击下等错误率超38%,实现强身份隐匿
  • 兼容主流语音模型,兼顾隐私性、多样性与语音质量

语音合成与语音转换技术的进步对个人隐私构成严重威胁,亟需强有力的说话人匿名系统(SAS)。现有方法在手工特征空间或说话人嵌入空间修改语音特征,往往难以生成足够多样的声音。本文提出NouveauVoice,一种基于层次化深度变分自编码器(NVAE)的新型伪说话人生成框架。作为独立插件模块,可集成至FACodec和CosyVoice2等先进架构中,利用可解析采样与证据下界(ELBO)目标,生成高度表达且多样性显著提升的伪说话人嵌入。在类语音隐私挑战赛的评测协议下,结合最大均值差异(MMD)分析,结果表明该方法在对抗自动说话人验证攻击时,等错误率(EER)超过38%,实现严格匿名性、丰富伪说话人多样性与下游语音实用性(如可懂度与情感表现力)之间的合理权衡。

原文摘要 · Abstract (English)

Advanced neural technologies in speech synthesis and voice conversion (VC) have introduced severe risks to personal privacy, necessitating robust Speaker Anonymization Systems (SAS). Existing SAS approaches modify voice characteristics in the hand-crafted feature space or speaker embedding space, often struggling to provide sufficient identity variance across generated voices. In this paper, we propose NouveauVoice, a novel pseudo-speaker generation framework based on a Hierarchical Deep Variational Autoencoder (NVAE). Integrated as a standalone plug-in module on top of state-of-the-art architectures (FACodec and CosyVoice2), our approach leverages tractable sampling and the Evidence Lower Bound (ELBO) objective to synthesize highly expressive pseudo-speaker embeddings with significantly enhanced speaker diversity. Evaluating our framework under a protocol similar to the VoicePrivacy Challenge alongside Maximum Mean Discrepancy (MMD) analysis, we demonstrate that NouveauVoice achieves strong identity concealment, yielding an Equal Error Rate (EER) exceeding 38% against an automatic speaker verification attacker model. Our system shows a reasonable trade-off between strict anonymity, rich pseudo-speaker diversity, and downstream speech utility, such as intelligibility and emotional expressiveness.

语音隐私说话人匿名生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。