arXiv:2604.17000eess.AS2026-04

保护语音隐私同时保留数据可用性,提升语音模型训练效果。

Anonymization, Not Elimination: Utility-Preserved Speech Anonymization

论文配图:Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
图 1 · 摘自论文原文
  • 两阶段框架:生成编辑替换敏感信息,流匹配生成多样匿名语音
  • 在ASR/TTS/SER任务中保持高精度,隐私保护优于基线方法
  • 新评估方案兼顾声学与内容隐私,更真实反映数据实用性

大规模语音数据的使用日益普遍,隐私保护成为关键问题。现有匿名化方法常导致语音连续性破坏或声纹多样性降低,影响自动语音识别(ASR)、文本转语音(TTS)和语音情感识别(SER)等下游任务的数据价值。当前评估多依赖预训练模型直接测试,视角有限。为此,我们提出一种新型两阶段框架,在保护语言内容和声纹身份的同时维持数据可用性。内容隐私通过生成式语音编辑模型无缝替换个人可识别信息;声纹隐私采用基于流匹配的F3-VA框架,三阶段设计生成多样且独特的匿名说话人。为实现更全面评估,我们结合声学与内容双维度的说话人验证指标衡量隐私保护,并从头训练ASR、TTS、SER模型评估实用性。实验表明,相比语音隐私挑战赛基线方法,本框架在显著增强隐私保护的同时,仅造成极小的性能损失,所提评估协议更真实反映了隐私保护下的数据实用性。

原文摘要 · Abstract (English)

The growing reliance on large-scale speech data has made privacy protection a critical concern. However, existing anonymization approaches often degrade data utility, for example by disrupting acoustic continuity or reducing vocal diversity, which compromises the value of speech data for downstream tasks such as Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and Speech Emotion Recognition (SER). Current evaluation practices are also limited, as they mainly rely on direct testing of anonymized speech with pretrained models, providing only a partial view of utility. To address these issues, we propose a novel two-stage framework that protects both linguistic content and acoustic identity while maintaining usability. For content privacy, we employ a generative speech editing model to seamlessly replace personally identifiable information (PII), and for voice privacy, we introduce F3-VA, a flow-matching-based anonymization framework with a three-stage design that produces diverse and distinct anonymized speakers. To enable a more comprehensive assessment, we evaluate privacy using both acoustic- and content-based speaker verification metrics, and assess utility by training ASR, TTS, and SER models from scratch. Experimental results show that our framework achieves stronger privacy protection with minimal utility degradation compared to baselines from the VoicePrivacy Challenge, while the proposed evaluation protocol provides a more realistic reflection of the utility of anonymized speech under privacy protection.

语音匿名化隐私保护数据可用性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。