用低维向量控制语音印象,支持自然语言直接生成目标音色。
Voice Impression Control in Zero-Shot TTS
- 用低维向量表示语音印象强度(如暗-亮)
- 主观与客观评估均验证了控制效果
- 可由大模型根据自然语言描述自动生成
语音中的非语言信息对听者印象的形成至关重要。尽管零样本文本到语音(TTS)已实现高说话人保真度,但调节细微的非语言信息以控制感知到的语音特征(即印象)仍具挑战。为此,我们提出一种零样本TTS中的语音印象控制方法,利用低维向量表示多种语音印象对(如暗-亮)的强度。客观与主观评估结果均证明该方法在印象控制上的有效性。此外,通过大语言模型生成该向量,可基于目标印象的自然语言描述实现自动化的指定印象生成,无需手动调优。音频示例可在演示页面获取:https://ntt-hilab-gensp.github.io/is2025voiceimpression/
原文摘要 · Abstract (English)
Para-/non-linguistic information in speech is pivotal in shaping the listeners' impression. Although zero-shot text-to-speech (TTS) has achieved high speaker fidelity, modulating subtle para-/non-linguistic information to control perceived voice characteristics, i.e., impressions, remains challenging. We have therefore developed a voice impression control method in zero-shot TTS that utilizes a low-dimensional vector to represent the intensities of various voice impression pairs (e.g., dark-bright). The results of both objective and subjective evaluations have demonstrated our method's effectiveness in impression control. Furthermore, generating this vector via a large language model enables target-impression generation from a natural language description of the desired impression, thus eliminating the need for manual optimization. Audio examples are available on our demo page (https://ntt-hilab-gensp.github.io/is2025voiceimpression/).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。