用音频嵌入序列相似度评估环境音合成质量,相关性优于传统方法。
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
- 基于音频嵌入序列的相似性计算,结合max-norm与p-norm。
- 在主观评分上相关性达0.82,显著高于传统指标。
- 适合需要高效评估环境音合成效果的研究者使用。
我们提出一种新型客观评价指标AudioBERTScore,用于文本到音频(TTA)中合成音频的评估。由于主观评价成本高,现有方法多依赖梅尔倒谱失真等指标,但其与主观评分的相关性较弱。AudioBERTScore通过计算合成音频与参考音频嵌入序列之间的相似性进行评估,不仅采用传统BERTScore中的max-norm,还引入p-norm以捕捉环境音的非局部特性。实验结果表明,该方法得到的分数与主观评价值的相关性高达0.82,显著优于现有指标。
原文摘要 · Abstract (English)
We propose a novel objective evaluation metric for synthesized audio in text-to-audio (TTA), aiming to improve the performance of TTA models. In TTA, subjective evaluation of the synthesized sound is an important, but its implementation requires monetary costs. Therefore, objective evaluation such as mel-cepstral distortion are used, but the correlation between these objective metrics and subjective evaluation values is weak. Our proposed objective evaluation metric, AudioBERTScore, calculates the similarity between embedding of the synthesized and reference sounds. The method is based not only on the max-norm used in conventional BERTScore but also on the $p$-norm to reflect the non-local nature of environmental sounds. Experimental results show that scores obtained by the proposed method have a higher correlation with subjective evaluation values than conventional metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。