arXiv:2603.27001eess.AScs.CL2026-03

实时语音匿名化系统,让非母语口音变母语口音,提升隐私保护效果。

PHONOS: PHOnetic Neutralization for Online Streaming Applications

  • 用静音感知的DTW对齐+零样本语音转换生成标准口音语音作为参考
  • 因果性口音转换模块仅需40毫秒前瞻,实现低延迟实时处理
  • 口音识别置信度下降81%,语音嵌入空间中说话人关联性减弱

说话人匿名化(SA)系统在保留地域或非母语口音的同时改变音色,但口音会缩小匿名范围。为此,我们提出PHONOS,一个用于实时流式应用的语音匿名化模块,可将非母语口音中和为近似母语发音。该方法预先生成黄金说话人语音,保留源音色与节奏,通过静音感知的动态时间规整(DTW)对齐与零样本语音转换,替换外源音段为母语音段。这些语音作为监督信号,训练一个因果性口音转换器,将非母语内容音素映射为母语对应项,最大前瞻仅40毫秒,采用联合交叉熵与连接时序分类(CTC)损失进行训练。评估显示,非母语口音识别置信度降低81%,听觉测试评分与此一致,且说话人关联性下降;同时在单张GPU上延迟低于241毫秒。

原文摘要 · Abstract (English)

Speaker anonymization (SA) systems modify timbre while leaving regional or non-native accents intact, which is problematic because accents can narrow the anonymity set. To address this issue, we present PHONOS, a streaming module for real-time SA that neutralizes non-native accent to sound native-like. Our approach pre-generates golden speaker utterances that preserve source timbre and rhythm but replace foreign segmentals with native ones using silence-aware DTW alignment and zero-shot voice conversion. These utterances supervise a causal accent translator that maps non-native content tokens to native equivalents with at most 40ms look-ahead, trained using joint cross-entropy and CTC losses. Our evaluations show an 81% reduction in non-native accent confidence, with listening-test ratings consistent with this shift, and reduced speaker linkability as accent-neutralized utterances move away from the original speaker in embedding space while having latency under 241 ms on single GPU.

语音匿名口音中和实时处理零样本转换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。