arXiv:2506.16310cs.LGcs.HC2025-06被引 1

让多语言语音合成同时精准模拟印地语口音与情感,支持实时口音切换。

Optimizing Multilingual Text-To-Speech with Accents & Emotions

  • 采用分层编码器解码器架构与文化敏感情感嵌入层,实现口音与情感解耦。
  • 口音准确率提升23.7%(WER从15.4%降至11.8%),情感识别率达85.3%。
  • 适合南亚教育科技与无障碍软件,支持真实场景下的跨语言语音生成。

当前最先进的文本转语音(TTS)系统在单语环境下已实现高自然度,但在多语言场景中仍难以准确合成印地语等语言的口音及符合语境的情感表达,主要因现有框架缺乏文化细微差异建模。本文提出一种新TTS架构,集成口音与音译保留能力,并引入多尺度情感建模,特别针对印地语和印度英语口音优化。方法上扩展Parler-TTS模型,采用语言特异性音素对齐的混合编码器-解码器架构,结合基于母语者语料训练的文化敏感情感嵌入层,并引入动态口音代码切换与残差向量量化机制。定量测试显示,口音准确率提升23.7%(词错误率从15.4%降至11.8%),母语听者情感识别率达85.3%,优于METTS与VECL-TTS基线。系统创新在于可实时混合语码生成,例如“Namaste, let's talk about <Hindi phrase>”,实现无缝口音切换且情感一致。主观评估200名用户反馈平均意见得分(MOS)达4.2/5,文化正确性显著优于现有系统(p<0.01)。本研究通过展示可扩展的口音-情感解耦范式,推动跨语言语音合成在南亚教育科技与无障碍软件中的应用。

原文摘要 · Abstract (English)

State-of-the-art text-to-speech (TTS) systems realize high naturalness in monolingual environments, synthesizing speech with correct multilingual accents (especially for Indic languages) and context-relevant emotions still poses difficulty owing to cultural nuance discrepancies in current frameworks. This paper introduces a new TTS architecture integrating accent along with preserving transliteration with multi-scale emotion modelling, in particularly tuned for Hindi and Indian English accent. Our approach extends the Parler-TTS model by integrating A language-specific phoneme alignment hybrid encoder-decoder architecture, and culture-sensitive emotion embedding layers trained on native speaker corpora, as well as incorporating a dynamic accent code switching with residual vector quantization. Quantitative tests demonstrate 23.7% improvement in accent accuracy (Word Error Rate reduction from 15.4% to 11.8%) and 85.3% emotion recognition accuracy from native listeners, surpassing METTS and VECL-TTS baselines. The novelty of the system is that it can mix code in real time - generating statements such as "Namaste, let's talk about <Hindi phrase>" with uninterrupted accent shifts while preserving emotional consistency. Subjective evaluation with 200 users reported a mean opinion score (MOS) of 4.2/5 for cultural correctness, much better than existing multilingual systems (p<0.01). This research makes cross-lingual synthesis more feasible by showcasing scalable accent-emotion disentanglement, with direct application in South Asian EdTech and accessibility software.

语音合成多语言口音建模情感识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。