arXiv:2603.14275eess.AScs.AI2026-03中稿 · Interspeech 2026 a…被引 1

通过离散扩散模型实现可调节口音强度的语音口音归一化。

Controllable Accent Normalization via Discrete Diffusion

  • 基于自监督语音标记的掩码离散扩散,选择性重用源语音标记以控制口音强度。
  • 在多口音英语数据上达到最低词错误率,同时保持自然语速和可解释控制。
  • 适合语言学习、配音等需灵活调节口音强度的应用场景。

现有口音归一化方法通常无法控制口音强度,而语言学习和配音等应用需要可调节的口音保留能力。本文提出 DLM-AN,一个基于自监督语音标记的掩码离散扩散的可控口音归一化系统。一个通用标记预测器识别可能编码母语发音的源标记;这些标记被选择性地用于初始化反向扩散过程,从而提供一种简单有效的口音强度控制机制:重用更多标记可保留更多原始口音。DLM-AN 进一步引入基于流匹配的时长比例预测器,自动调整总时长以更好地匹配母语节奏。在多口音英语数据上的实验表明,DLM-AN 在所有对比系统中实现了最低的词错误率,同时表现出优异的口音降低效果和平滑、可解释的口音强度控制能力。

原文摘要 · Abstract (English)

Existing accent normalization methods do not typically offer control over accent strength, yet many applications-such as language learning and dubbing-require tunable accent retention. We propose DLM-AN, a controllable accent normalization system built on masked discrete diffusion over self-supervised speech tokens. A Common Token Predictor identifies source tokens that likely encode native pronunciation; these tokens are selectively reused to initialize the reverse diffusion process. This provides a simple yet effective mechanism for controlling accent strength: reusing more tokens preserves more of the original accent. DLM-AN further incorporates a flow-matching Duration Ratio Predictor that automatically adjusts the total duration to better match the native rhythm. Experiments on multi-accent English data show that DLM-AN achieves the lowest word error rate among all compared systems while delivering competitive accent reduction and smooth, interpretable accent strength control.

语音处理扩散模型口音归一化可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。