提出动态音色表征,实现低延迟语音转换与匿名化。
TVTSyn: Content-Synchronous Time-Varying Timbre for Streaming Voice Conversion and Anonymization
- 用随内容变化的音色表示,对齐语义与身份的时间粒度。
- 端到端流式处理,GPU延迟低于80毫秒,自然度与匿名性提升。
- 适合实时语音隐私保护与高保真语音合成场景。
实时语音转换与说话人匿名化要求因果、低延迟的合成,同时保持可懂性和自然度。现有系统存在核心表征错配:内容是时变的,而说话人身份以静态全局嵌入注入。本文提出一种可流式处理的语音合成器,通过内容同步的时变音色(TVT)表示对齐身份与内容的时间粒度。全局音色记忆将全局音色实例扩展为多个紧凑分量;帧级内容关注该记忆,门控调节变化,球面插值在保持身份几何的同时实现平滑局部变化。此外,因子化向量量化瓶颈正则化内容,减少残留说话人泄漏。所提系统为端到端可流式处理,GPU延迟低于80毫秒。实验表明,在自然度、说话人迁移和匿名化方面优于当前最优流式基线,确立了TVT在严格延迟预算下的可扩展性,适用于隐私保护与富表达语音合成。
原文摘要 · Abstract (English)
Real-time voice conversion and speaker anonymization require causal, low-latency synthesis without sacrificing intelligibility or naturalness. Current systems have a core representational mismatch: content is time-varying, while speaker identity is injected as a static global embedding. We introduce a streamable speech synthesizer that aligns the temporal granularity of identity and content via a content-synchronous, time-varying timbre (TVT) representation. A Global Timbre Memory expands a global timbre instance into multiple compact facets; frame-level content attends to this memory, a gate regulates variation, and spherical interpolation preserves identity geometry while enabling smooth local changes. In addition, a factorized vector-quantized bottleneck regularizes content to reduce residual speaker leakage. The resulting system is streamable end-to-end, with <80 ms GPU latency. Experiments show improvements in naturalness, speaker transfer, and anonymization compared to SOTA streaming baselines, establishing TVT as a scalable approach for privacy-preserving and expressive speech synthesis under strict latency budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。