质疑情感嵌入相似性评估的有效性,发现其易被语音模仿误导。
The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation
- 用对抗任务测试情感嵌入相似性,检验其是否真正捕捉情绪特征。
- 高分类准确率下仍无法有效区分不同情绪,与人类感知不一致。
- 适合关注语音合成评估方法可靠性的研究者和工程师。
情感表达的客观评估对语音生成至关重要,尤其在需要情感韵律迁移的风格化合成与语音转换中。当前普遍采用参考样本与生成样本之间的情感嵌入相似性(如emotion2vec)的余弦相似度来量化情感表现力,该方法假设嵌入能捕捉情感线索,不受语言和说话人差异影响。本文通过受控的对抗任务与人类感知对齐测试挑战这一假设。尽管分类准确率很高,但这些潜在空间因表征局限,导致语言和说话人干扰掩盖了情感特征,削弱了判别能力,使指标与人类感知严重脱节。该声学脆弱性表明,现有方法更奖励声学模仿而非真实情感合成。
原文摘要 · Abstract (English)
Objective metrics for emotional expressiveness are vital for speech generation, particularly in expressive synthesis and voice conversion requiring emotional prosody transfer. To quantify this, the field widely relies on emotion similarity between reference and generated samples. This approach computes cosine similarity of embeddings from encoders like emotion2vec, assuming they capture affective cues despite linguistic and speaker variations. We challenge this assumption through controlled adversarial tasks and human alignment tests. Despite high classification accuracy, these latent spaces are unsuitable for zero-shot similarity evaluation. Representational limitations cause linguistic and speaker interference to overshadow emotional features, degrading discriminative ability. Consequently, the metric misaligns with human perception. This acoustic vulnerability reveals it rewards acoustic mimicry over genuine emotional synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。