无需带口音数据,就能自由控制多语言语音合成的口音。
Accent Vector: Controllable Accent Manipulation for Multilingual TTS Without Accented Data
- 通过非英语母语语音微调模型,提取口音特征向量。
- 可调节向量强度,实现口音强弱与混合口音的精细控制。
- 支持多语言口音操控,适合跨语言语音应用开发。
口音是社会文化的重要体现,塑造个体身份表达。多数英语使用者为非母语者(L2),但当前文本转语音(TTS)系统主要依赖美式口音数据,受限于带口音训练数据。本文提出「Accent Vector」,一种无需显式口音数据即可实现多语言TTS中口音可控调整的表示方法。该向量通过在另一种语言(非英语)的母语语音上微调TTS模型,并计算其捕捉英语口音特征的任务向量得到。通过缩放和插值向量,可实现口音强度的精细调节及混合口音语音生成。此外,该方法具备跨语言泛化能力,支持多种语言的口音控制。客观与人工评估均证实其在细粒度与组合性口音控制上的有效性。
原文摘要 · Abstract (English)
Accent is an integral part of society, reflecting multiculturalism and shaping how individuals express identity. The majority of English speakers are non-native (L2) speakers, yet current Text-To-Speech (TTS) systems primarily model American-accented English due limited accented data. We propose \textit{Accent Vector}, a controllable representation that enables accent manipulation in multilingual TTS without requiring accented training data. \textit{Accent Vector} is derived by fine-tuning a TTS system on native speech of a different language (i.e. non-English) and computing task vectors capturing accent characteristics (i.e. in English). By scaling and interpolating the vector, we achieve fine-grained control over accent strength and generate mixed-accent speech. In addition, it generalizes beyond English, enabling accent control across multiple languages. Objective and human evaluations confirm the effectiveness of Accent Vector for fine-grained and compositional accent control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。