arXiv:2606.20218cs.SD2026-06中稿 · Interspeech 2026

用说话人匿名化实现零延迟语音转换,兼顾音色隐藏与语调保真。

Zero-VC: Zero-Lookahead Streaming Voice Conversion via Speaker Anonymization

  • 引入说话人匿名化作为新扰动机制,分离音色与语言内容。
  • 无需缓冲未来帧,实现严格因果推理,延迟归零。
  • 适合对实时性要求高的语音转换场景,如直播、通话系统。

流式零样本语音转换在不降低语用价值或增加延迟的前提下,难以将音色与语言内容解耦。现有方法依赖信息瓶颈(IB)或说话人扰动,但IB会丢弃韵律特征,迫使模型显式注入基频等特征,常需缓冲未来帧,导致算法前瞻延迟。而现有扰动方法未充分考虑音色泄漏与语用保真之间的权衡。本文发现说话人匿名化(SA)的内在目标恰好契合这一平衡。因此,我们提出将SA作为新型扰动机制,在有效抑制音色泄漏的同时保留韵律信息。关键在于,SA生成的鲁棒表征显著降低了生成器对后续上下文的依赖,从而实现严格因果、零前瞻的网络结构。音频样例可访问 https://amphionteam.github.io/Zero-VC-demo/。

原文摘要 · Abstract (English)

Streaming zero-shot voice conversion struggles to disentangle timbre from linguistic content without degrading utility or inflating latency. Current methods rely on information bottleneck (IB) or speaker perturbation. While IB filters out timbre, it discards prosody, forcing models to explicitly inject features like fundamental frequency. This often requires buffering future frames, creating algorithmic lookahead latency. On the other hand, existing perturbation methods largely overlook the crucial trade-off between timbre leakage and utility preservation. Recognizing this neglected trade-off, we find that the inherent objective of Speaker Anonymization (SA) aligns well with balancing these factors. Thus, we introduce SA as a novel perturbation mechanism to explicitly mitigate timbre leakage while retaining prosodic utility. Crucially, SA's robust representations significantly alleviate the generator's reliance on future context, enabling our strictly causal, zero-lookahead network. Audio samples are available at https://amphionteam.github.io/Zero-VC-demo/.

语音转换零延迟说话人匿名

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。