arXiv:2603.07550cs.CLcs.AI2026-03被引 2

无需训练数据,用音系规则生成特定口音语音

Learning-free L2-Accented Speech Generation using Phonological Rules

  • 用音系规则在音素层面转换口音,不依赖口音语料
  • 实现西班牙语和印度口音英语的准确模拟,保持可懂性
  • 适合需要灵活控制口音的语音合成应用

口音在语音技术中对说话人身份和包容性至关重要。现有口音文本转语音(TTS)系统要么需要大规模口音数据集,要么缺乏音素级精细控制能力。本文提出一种结合音系规则与多语言TTS模型的口音化TTS框架。通过在音素序列上应用规则,实现音素层面的口音转换,同时保持语音可懂性。该方法无需任何口音训练数据,支持显式的音素级口音操控。我们为西班牙语和印度口音英语设计了规则集,建模了因音位约束导致的辅音、元音及音节结构系统性差异。分析了音素时长对齐与口音表现之间的权衡。实验表明,该方法能有效实现口音迁移,同时保持高质量语音输出。

原文摘要 · Abstract (English)

Accent plays a crucial role in speaker identity and inclusivity in speech technologies. Existing accented text-to-speech (TTS) systems either require large-scale accented datasets or lack fine-grained phoneme-level controllability. We propose a accented TTS framework that combines phonological rules with a multilingual TTS model. The rules are applied to phoneme sequences to transform accent at the phoneme level while preserving intelligibility. The method requires no accented training data and enables explicit phoneme-level accent manipulation. We design rule sets for Spanish- and Indian-accented English, modeling systematic differences in consonants, vowels, and syllable structure arising from phonotactic constraints. We analyze the trade-off between phoneme-level duration alignment and accent as realized in speech timing. Experimental results demonstrate effective accent shift while maintaining speech quality.

语音合成口音生成音系规则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。