arXiv:2507.17735eess.AScs.SD2025-07中稿 · INTERSPEECH 2025被引 5

用自监督离散标记实现无平行数据的口音标准化,保持说话人特征。

Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data

  • 通过自监督提取语音离散标记,跨口音转换
  • 在多英语口音上自然度与口音减少效果更优
  • 适合字幕配音等需时长保留的应用

口音标准化将外源口音语音转换为接近母语发音,同时保留说话人身份。本文提出一种新流程:利用自监督离散标记与非平行训练数据,从源语音中提取标记,经专用模型转换后,通过流匹配合成输出。该方法在自然度、口音减弱和音色保留方面均优于帧对帧基线,在多种英语口音下表现优异。通过标记级语音分析验证了基于标记方法的有效性。此外,开发了两种时长保留方法,适用于字幕配音等场景。

原文摘要 · Abstract (English)

Accent normalization converts foreign-accented speech into native-like speech while preserving speaker identity. We propose a novel pipeline using self-supervised discrete tokens and non-parallel training data. The system extracts tokens from source speech, converts them through a dedicated model, and synthesizes the output using flow matching. Our method demonstrates superior performance over a frame-to-frame baseline in naturalness, accentedness reduction, and timbre preservation across multiple English accents. Through token-level phonetic analysis, we validate the effectiveness of our token-based approach. We also develop two duration preservation methods, suitable for applications such as dubbing.

口音转换自监督语音合成非平行数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。