arXiv:2603.05977eess.AS2026-03被引 1

不训练模型,推理时就能消除语音口音并保留原嗓音。

Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech

  • 通过分析模型内部激活差异,提取分层的引导向量。
  • 推理时应用向量,使输出语音口音中性化且保留原嗓音特征。
  • 无需重新训练,对未见口音说话人也有效,适合语音克隆场景。

零样本文本转语音模型能生成兼具参考说话人音色和口音的语音,但音色与口音常难以分离。本文提出一种新颖的、后处理的、无需训练的方法,利用推理时的激活引导实现口音中性化,同时保留说话人原始音色。首先离线提取模型各层中因口音差异产生的“引导向量”,这些向量基于模型内部激活差异构建。推理时,将这些向量注入模型,引导生成口音中性、音色保持的语音。实验表明,该方法能有效降低输出口音,且在未见过的口音说话人上具有强泛化能力,为无口音语音克隆提供了实用方案。

原文摘要 · Abstract (English)

Zero-shot Text-to-Speech (TTS) models can generate speech that captures both the voice timbre and accent of a reference speaker. However, disentangling these attributes remains challenging, as the output often inherits both the accent and timbre from the reference. In this study, we introduce a novel, post-hoc, and training-free approach to neutralize accent while preserving the speaker's original timbre, utilizing inference-time activation steering. We first extract layer-specific "steering vectors" offline, which are derived from the internal activation differences within the TTS model between accented and native speech. During inference, the steering vectors are applied to guide the model to produce accent-neutralized, timbre-preserving speech. Empirical results demonstrate that the proposed steering vectors effectively mitigate the output accent and exhibit strong generalizability to unseen accented speakers, offering a practical solution for accent-free voice cloning.

语音合成零样本口音中性化推理优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。