改进音频-文本匹配评估,让模型更贴近人类感知
Human-CLAP: Human-perception-based contrastive language-audio pretraining
- 用人类主观评分训练新的对比语言音频模型
- 提升评估相关性,比传统方法高出0.25以上
- 适合需要真实用户感知的音频生成与评价场景
对比语言音频预训练(CLAP)广泛应用于音频生成与识别任务。例如,基于CLAP嵌入相似性的CLAPScore已成为文本到音频任务中评估音频与文本相关性的主要指标。然而,CLAPScore与人类主观评价分数之间的关系尚不明确。我们发现,CLAPScore与人类主观评分的相关性较低。为此,提出一种基于人类感知的CLAP模型——Human-CLAP,通过使用主观评分数据训练对比语言音频模型。实验结果表明,相较于传统CLAP,Human-CLAP使CLAPScore与主观评价分数之间的斯皮尔曼等级相关系数(SRCC)提升了超过0.25。
原文摘要 · Abstract (English)
Contrastive language-audio pretraining (CLAP) is widely used for audio generation and recognition tasks. For example, CLAPScore, which utilizes the similarity of CLAP embeddings, has been a major metric for the evaluation of the relevance between audio and text in text-to-audio. However, the relationship between CLAPScore and human subjective evaluation scores is still unclarified. We show that CLAPScore has a low correlation with human subjective evaluation scores. Additionally, we propose a human-perception-based CLAP called Human-CLAP by training a contrastive language-audio model using the subjective evaluation score. In our experiments, the results indicate that our Human-CLAP improved the Spearman's rank correlation coefficient (SRCC) between the CLAPScore and the subjective evaluation scores by more than 0.25 compared with the conventional CLAP.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。