arXiv:2506.23582cs.SDeess.AS2025-06中稿 · INTERSPEECH2025被引 7

构建首个文本音频相关性主观评价数据集,支持自动评估模型优化。

RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio

  • 构建RELATE数据集,人工标注文本与音频的相关性评分。
  • 新模型在多个音效类别上超越CLAPScore,预测准确率更高。
  • 适合语音生成、多模态评估研究者使用。

在文本到音频(TTA)研究中,输入文本与输出音频的相关性是重要评估维度。传统方法依赖主观与客观双重视角,但主观评估成本高,而客观评估与主观评分的相关性不明确。本文构建了开源的RELATE数据集,专门用于主观评估文本与音频的相关性,并基准测试了一个从合成音频中自动预测主观评分的模型。该模型性能优于传统的CLAPScore,在多个声音类别中均表现更优,验证了其广泛适用性。

原文摘要 · Abstract (English)

In text-to-audio (TTA) research, the relevance between input text and output audio is an important evaluation aspect. Traditionally, it has been evaluated from both subjective and objective perspectives. However, subjective evaluation is costly in terms of money and time, and objective evaluation is unclear regarding the correlation to subjective evaluation scores. In this study, we construct RELATE, an open-sourced dataset that subjectively evaluates the relevance. Also, we benchmark a model for automatically predicting the subjective evaluation score from synthesized audio. Our model outperforms a conventional CLAPScore model, and that trend extends to many sound categories.

文本音频主观评价数据集多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。