让语音匿名时保留情绪,实时无延迟。
StreamVoiceAnon+: Emotion-Preserving Streaming Speaker Anonymization via Frame-Level Acoustic Distillation
- 用同一说话人中性语句微调,帧级情感蒸馏保情绪
- 情绪保留率49.2%(+24%相对提升),可懂度5.77% WER
- 适合需要实时隐私保护且重情绪的语音应用
针对流式语音匿名化中情绪丢失问题,现有神经音频编解码语言模型在音频续写时易丢失情感信息,因内容标记舍弃情感特征,模型趋向默认声学模式而非保留副语言属性。本文提出基于同说话人中性语句对的监督微调,并在声学标记隐藏状态上实施帧级情感蒸馏。所有改进仅限微调阶段,4块GPU耗时不足2小时,推理零延迟开销,仍保持180ms流式延迟。在VoicePrivacy 2024协议下,本方法实现49.2% UAR(情绪保留率)与5.77% WER(可懂度),相比基线39.7%提升24%,优于情感提示变体(44.6% UAR),同时保障强隐私性(EER 49.0%)。演示与代码已公开:https://anonymous3842031239.github.io/
原文摘要 · Abstract (English)
We address the challenge of preserving emotional content in streaming speaker anonymization (SA). Neural audio codec language models trained for audio continuation tend to degrade source emotion: content tokens discard emotional information, and the model defaults to dominant acoustic patterns rather than preserving paralinguistic attributes. We propose supervised finetuning with neutral-emotion utterance pairs from the same speaker, combined with frame-level emotion distillation on acoustic token hidden states. All modifications are confined to finetuning, which takes less than 2 hours on 4 GPUs and adds zero inference latency overhead, while maintaining a competitive 180ms streaming latency. On the VoicePrivacy 2024 protocol, our approach achieves a 49.2% UAR (emotion preservation) with 5.77% WER (intelligibility), a +24% relative UAR improvement over the baseline (39.7%->49.2%) and +10% over the emotion-prompt variant (44.6% UAR), while maintaining strong privacy (EER 49.0%). Demo and code are available: https://anonymous3842031239.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。