arXiv:2602.20113cs.SDcs.AI2026-02被引 1

首个实时零样本语音风格转换系统,可秒级完成音色情感迁移。

StyleStream: Real-Time Zero-Shot Voice Style Conversion

  • 采用去风格化与风格重建双模块,通过文本监督和信息瓶颈实现内容与风格解耦
  • 端到端延迟仅1秒,支持非自回归实时处理,性能达当前最优
  • 适合语音克隆、虚拟主播等需快速适配声音风格的场景

语音风格转换旨在将输入语音转换为目标说话人的音色、口音和情感,核心挑战在于内容与风格的解耦。现有方法转化质量有限,且未解决实时性问题。本文提出StyleStream,首个可流式处理的零样本语音风格转换系统,达到当前最佳性能。系统包含两个组件:去风格化器(Destylizer)在保留语言内容的同时去除风格特征;风格重建器(Stylizer)为扩散变换器(DiT),根据参考语音条件重建目标风格。通过文本监督与强约束信息瓶颈,实现鲁棒的内容-风格解耦,设计为全非自回归架构,端到端延迟仅1秒,支持实时转换。演示与样例见:https://berkeley-speech-group.github.io/StyleStream/

原文摘要 · Abstract (English)

Voice style conversion aims to transform an input utterance to match a target speaker's timbre, accent, and emotion, with a central challenge being the disentanglement of linguistic content from style. While prior work has explored this problem, conversion quality remains limited, and real-time voice style conversion has not been addressed. We propose StyleStream, the first streamable zero-shot voice style conversion system that achieves state-of-the-art performance. StyleStream consists of two components: a Destylizer, which removes style attributes while preserving linguistic content, and a Stylizer, a diffusion transformer (DiT) that reintroduces target style conditioned on reference speech. Robust content-style disentanglement is enforced through text supervision and a highly constrained information bottleneck. This design enables a fully non-autoregressive architecture, achieving real-time voice style conversion with an end-to-end latency of 1 second. Samples and real-time demo: https://berkeley-speech-group.github.io/StyleStream/.

语音转换实时处理零样本扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。