用自监督离散标记实现无平行数据的口音标准化,保持说话人特征。
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
- 通过自监督提取语音离散标记,跨口音转换
- 在多英语口音上自然度与口音减少效果更优
- 适合字幕配音等需时长保留的应用
口音标准化将外源口音语音转换为接近母语发音,同时保留说话人身份。本文提出一种新流程:利用自监督离散标记与非平行训练数据,从源语音中提取标记,经专用模型转换后,通过流匹配合成输出。该方法在自然度、口音减弱和音色保留方面均优于帧对帧基线,在多种英语口音下表现优异。通过标记级语音分析验证了基于标记方法的有效性。此外,开发了两种时长保留方法,适用于字幕配音等场景。
原文摘要 · Abstract (English)
Accent normalization converts foreign-accented speech into native-like speech while preserving speaker identity. We propose a novel pipeline using self-supervised discrete tokens and non-parallel training data. The system extracts tokens from source speech, converts them through a dedicated model, and synthesizes the output using flow matching. Our method demonstrates superior performance over a frame-to-frame baseline in naturalness, accentedness reduction, and timbre preservation across multiple English accents. Through token-level phonetic analysis, we validate the effectiveness of our token-based approach. We also develop two duration preservation methods, suitable for applications such as dubbing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。