用少量标注数据构建语音表达力客观评分框架
Decoding the Ear: A Framework for Objectifying Expressiveness from Human Preference Through Efficient Alignment
- 基于语音学与心理学,从情感、语调、自然度三维度量化表达力
- 仅需500样本即达0.86相关性,显著优于传统评估方法
- 适合语音合成优化与高质量数据筛选,支持模型迭代
近期语音到语音(S2S)模型虽能生成清晰语音,但缺乏自然表达力,主因是缺乏可靠评估指标。现有方法如主观MOS评分、低级声学特征和情绪识别存在成本高、覆盖不全等问题。为此,我们提出DeEAR(Decoding the Expressive Preference of eAR)框架,将人类对语音表达力的偏好转化为客观分数。该框架基于语音学与心理学,从情感、语调和自然度三个维度评估语音,在少于500个标注样本下实现与人类感知高度一致(斯皮尔曼等级相关系数,SRCC = 0.86)。除提供可靠评分外,DeEAR还支持公平基准测试与定向数据筛选,不仅能区分不同S2S模型的表达力差距,还能从中选出14,000条表达力强的语音,构建ExpressiveSpeech数据集,使S2S模型表达力得分从2.0提升至23.4(满分100分)。演示与代码已公开于https://github.com/FreedomIntelligence/ExpressiveSpeech。
原文摘要 · Abstract (English)
Recent speech-to-speech (S2S) models generate intelligible speech but still lack natural expressiveness, largely due to the absence of a reliable evaluation metric. Existing approaches, such as subjective MOS ratings, low-level acoustic features, and emotion recognition are costly, limited, or incomplete. To address this, we present DeEAR (Decoding the Expressive Preference of eAR), a framework that converts human preference for speech expressiveness into an objective score. Grounded in phonetics and psychology, DeEAR evaluates speech across three dimensions: Emotion, Prosody, and Spontaneity, achieving strong alignment with human perception (Spearman's Rank Correlation Coefficient, SRCC = 0.86) using fewer than 500 annotated samples. Beyond reliable scoring, DeEAR enables fair benchmarking and targeted data curation. It not only distinguishes expressiveness gaps across S2S models but also selects 14K expressive utterances to form ExpressiveSpeech, which improves the expressive score (from 2.0 to 23.4 on a 100-point scale) of S2S models. Demos and codes are available at https://github.com/FreedomIntelligence/ExpressiveSpeech
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。