arXiv:2507.08012cs.CLcs.AI2025-07

通过重复微调发现语音模型隐含特征,提升可控性。

RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning

  • 用主成分分析挖掘合成样本的潜在变量作为新标签
  • 在冰岛语数据上实现连续与离散特征的可控提升
  • 适合语音合成控制、可解释性研究者使用

基于提示的文本转语音模型可通过自然语言指令控制语速、性别感知等语音特性。然而这类方法一方面受限:仅能控制训练时暴露的声学特征;另一方面又过于灵活:相同输入会产生不可控的变异,反映在语料统计中。本文提出一种新型微调策略,通过利用模型输出的不可控方差,对数千个合成样本进行主成分分析,识别出解释最大输出方差的潜在特征,并将其作为新标签用于二次微调。我们在两个基于表达性冰岛语语料库训练的模型上评估该方法,一个带情感披露,一个不带。对于无情感披露的模型,该方法成功提取出连续与离散特征,显著提升模型整体可控性。

原文摘要 · Abstract (English)

A Prompt-based Text-To-Speech model allows a user to control different aspects of speech, such as speaking rate and perceived gender, through natural language instruction. Although user-friendly, such approaches are on one hand constrained: control is limited to acoustic features exposed to the model during training, and too flexible on the other: the same inputs yields uncontrollable variation that are reflected in the corpus statistics. We investigate a novel fine-tuning regime to address both of these issues at the same time by exploiting the uncontrollable variance of the model. Through principal component analysis of thousands of synthesised samples, we determine latent features that account for the highest proportion of the output variance and incorporate them as new labels for secondary fine-tuning. We evaluate the proposed methods on two models trained on an expressive Icelandic speech corpus, one with emotional disclosure and one without. In the case of the model without emotional disclosure, the method yields both continuous and discrete features that improve overall controllability of the model.

语音合成特征发现可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。