用神经音频编解码器实现低延迟语音匿名,兼顾清晰度与情感保留。
Stream-Voice-Anon: Enhancing Utility of Real-Time Speaker Anonymization via Neural Audio Codec and Language Models
- 基于因果语言模型与量化内容码,分离说话人特征并注入伪身份信息
- 实时场景下语音识别错误率降低46%,情绪识别准确率提升28%
- 适合对隐私保护和语音质量要求高的实时语音应用
保护说话人身份对在线语音应用至关重要,但流式语音匿名(SA)仍研究不足。近期研究表明,神经音频编解码器(NAC)能有效分离说话人特征并保持语言保真度,且可与因果语言模型(LM)结合以增强语言保真度和提示控制。然而,现有基于NAC的在线LM系统专用于语音转换(VC),缺乏隐私保护技术。本文提出Stream-Voice-Anon,通过整合匿名化技术,将现代因果LM-based NAC架构适配至流式SA任务。其方法包括伪说话人表征采样、说话人嵌入混合及多样化提示选择策略,利用量化内容码的解耦特性防止说话人信息泄露。同时比较动态与固定延迟配置,探索实时场景下的延迟-隐私权衡。在VoicePrivacy 2024挑战协议下,相比前序最先进方法DarkStream,Stream-Voice-Anon实现语音识别错误率最高46%相对降低,情绪识别准确率最高28%相对提升,延迟维持在180ms(vs 200ms),对懒惰攻击者具备相当隐私保护能力,但对半知情攻击者有15%相对退化。
原文摘要 · Abstract (English)
Protecting speaker identity is crucial for online voice applications, yet streaming speaker anonymization (SA) remains underexplored. Recent research has demonstrated that neural audio codec (NAC) provides superior speaker feature disentanglement and linguistic fidelity. NAC can also be used with causal language models (LM) to enhance linguistic fidelity and prompt control for streaming tasks. However, existing NAC-based online LM systems are designed for voice conversion (VC) rather than anonymization, lacking the techniques required for privacy protection. Building on these advances, we present Stream-Voice-Anon, which adapts modern causal LM-based NAC architectures specifically for streaming SA by integrating anonymization techniques. Our anonymization approach incorporates pseudo-speaker representation sampling, a speaker embedding mixing and diverse prompt selection strategies for LM conditioning that leverage the disentanglement properties of quantized content codes to prevent speaker information leakage. Additionally, we compare dynamic and fixed delay configurations to explore latency-privacy trade-offs in real-time scenarios. Under the VoicePrivacy 2024 Challenge protocol, Stream-Voice-Anon achieves substantial improvements in intelligibility (up to 46% relative WER reduction) and emotion preservation (up to 28% UAR relative) compared to the previous state-of-the-art streaming method DarkStream while maintaining comparable latency (180ms vs 200ms) and privacy protection against lazy-informed attackers, though showing 15% relative degradation against semi-informed attackers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。