提出新评估指标SPAM,精准衡量语音合成对风格提示的遵循程度。
SPAM: Style Prompt Adherence Metric for Prompt-based TTS
- 将语音分解为声学特征,对齐风格提示,实现可解释评估
- 与人工评分相关性高(MOS),能区分不同风格语义
- 适合语音合成研究者用于自动评估风格一致性
基于提示的文本到语音(TTS)旨在生成符合文本提示中细微风格线索的语音。然而,以往方法缺乏可信且忠实的评估方式,无法保证评价是否真正基于提示或接近人类判断。为此,我们提出一种新的自动评估指标——风格提示遵循度度量(SPAM),明确满足合理性和忠实性。受CLAP启发,该方法将语音分解为声学属性,并将其与风格提示对齐;同时采用监督对比损失训练评分器,增强不同语义间的区分能力。我们在两个角度进行了实验:合理性实验表明,SPAM与平均意见分(MOS)具有强相关性;忠实性实验显示,SPAM能有效区分提示中的不同语义,证明其确实锚定于给定风格提示。我们认为,SPAM可为合成语音的风格提示遵循度提供可行的自动化评估方案。
原文摘要 · Abstract (English)
Prompt-based text-to-speech (TTS) aims to generate speech that adheres to fine-grained style cues provided in a text prompt. However, most prior works depend on neither plausible nor faithful measures to evaluate prompt adherence. That is, they cannot ensure whether the evaluation is grounded on the prompt and is similar to a human. Thus, we present a new automatic metric, the Style Prompt Adherence Metric, which explicitly satisfies both plausibility and faithfulness. Inspired by the CLAP, our approach factorizes speech into acoustic attributes and aligns them with the style prompt. Also, we trained the scorer with a supervised contrastive loss, which could provide a clearer distinction between different semantics. We conducted two experiments on two perspectives. The plausibility experiment showed that SPAM achieved a strong correlation with the mean opinion score (MOS). Also, the faithfulness experiment demonstrated that SPAM is successfully grounded to the given style prompt, as it can discriminate different semantics of the prompt. We believe that SPAM can provide a viable automatic solution for evaluating style prompt adherence of synthesized speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。