arXiv:2511.11104cs.SDcs.CL2025-11被引 2

解决语音合成中的口音与语言偏见,让不同方言发音更真实自然。

CLARITY: Contextual Linguistic Adaptation and Accent Retrieval for Dual-Bias Mitigation in Text-to-Speech Generation

  • 通过上下文语言适配,将文本精准匹配目标方言。
  • 引入检索增强口音提示,提升口音一致性与准确性。
  • 在12种英语口音上验证,显著改善语音感知质量。

指令引导的文本到语音(TTS)研究已达到高质量生成水平,但两种相互关联的偏差仍影响感知质量:口音偏差使模型偏向主流发音模式,语言偏差导致方言特有的词汇或文化信息错位。这两者共同影响真实口音生成。我们提出CLARITY(上下文语言适配与口音检索用于包容性语音合成),一个与主干模型无关的双信号优化框架。首先,采用上下文语言适配将输入文本本地化以匹配目标方言;其次,提出检索增强口音提示(RAAP),确保语音提示与口音一致。我们在12种英语口音上通过主观与客观分析评估了CLARITY,结果表明其显著提升了口音准确率与公平性,输出感知质量更高。代码与音频样本可于https://github.com/ICT-SIT/CLARITY 获取。

原文摘要 · Abstract (English)

Instruction-guided text-to-speech (TTS) research has reached a maturity level where excellent speech generation quality is possible on demand, yet two coupled biases persist in reducing perceived quality: accent bias, where models default towards dominant phonetic patterns, and linguistic bias, a misalignment in dialect-specific lexical or cultural information. These biases are interdependent and authentic accent generation requires both accent fidelity and correctly localized text. We present CLARITY (Contextual Linguistic Adaptation and Retrieval for Inclusive TTS sYnthesis), a backbone-agnostic framework to address both biases through dual-signal optimization. Firstly, we apply contextual linguistic adaptation to localize input text to align with the target dialect. Secondly, we propose retrieval-augmented accent prompting (RAAP) to ensure accent-consistent speech prompts. We evaluate CLARITY on twelve varieties of English accent via both subjective and objective analysis. Results clearly indicate that CLARITY improves accent accuracy and fairness, ensuring higher perceptual quality output\footnote{Code and audio samples are available at https://github.com/ICT-SIT/CLARITY.

语音合成口音生成公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。