arXiv:2510.10785cs.SD2025-10中稿 · ICASSP 2026被引 5

让语音改口音可调节,说话人身份更保真

FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec

  • 用分解式语音编码器分离口音与发音特征
  • 用户可调口音强度,转换效果媲美顶尖模型
  • 适合需要精细控制口音的语音合成应用

以往的口音转换方法缺乏对修改程度的显式控制。由于口音变化会影响说话人识别,平衡转换强度与身份保留至关重要。本文提出一种新的口音转换框架,提供显式的、用户可调控的参数以调节发音层面的口音修改强度。实验表明,该方法性能接近最新系统,显著提升说话人身份保留能力,并首次实现可控的零样本外语口音转换。

原文摘要 · Abstract (English)

Previous accent conversion (AC) methods, including foreign accent conversion (FAC), lack explicit control over the degree of modification. Because accent modification can alter the perceived speaker identity, balancing conversion strength and identity preservation is crucial. We present an AC framework that provides an explicit, user-controllable parameter to adjust the strength of pronunciation-level accent modification. Results show performance comparable to recent AC systems, stronger preservation of speaker identity, and unique support for controllable accent conversion.

语音转换口音迁移可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。