arXiv:2508.02295eess.AScs.SD2025-08被引 7

无需参考音源即可消除语音中的性别特征,保护隐私

Reference-free Adversarial Sex Obfuscation in Speech

  • 采用条件对抗学习分离语言内容与性别声学特征
  • 通过正则化使基频和共振峰轨迹符合中性性别特征
  • 在半知情攻击下仍优于现有方法,适合隐私敏感场景

语音性别转换存在数据收集带来的隐私风险,且即使无目标说话人参考,输出仍可能残留性别线索。我们提出 RASO(Reference-free Adversarial Sex Obfuscation),创新性地构建性别条件对抗学习框架,以解耦语言内容与性别相关声学标记,并引入显式正则化,使基频分布与共振峰轨迹对齐由性别平衡训练数据学习到的中性性别特征。RASO 能有效保留语言内容,在半知情攻击模型下显著优于现有方法。

原文摘要 · Abstract (English)

Sex conversion in speech involves privacy risks from data collection and often leaves residual sex-specific cues in outputs, even when target speaker references are unavailable. We introduce RASO for Reference-free Adversarial Sex Obfuscation. Innovations include a sex-conditional adversarial learning framework to disentangle linguistic content from sex-related acoustic markers and explicit regularisation to align fundamental frequency distributions and formant trajectories with sex-neutral characteristics learned from sex-balanced training data. RASO preserves linguistic content and, even when assessed under a semi-informed attack model, it significantly outperforms a competing approach to sex obfuscation.

语音隐私性别去标识对抗学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。