arXiv:2606.17404eess.AScs.SD2026-06中稿 · presentation at In…

通过声学事件级对齐,提升文本到音频生成的自动评估精度。

ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation

论文配图:ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation
图 1 · 摘自论文原文
  • 基于文本查询分解音频中的声学事件,实现细粒度对齐评估。
  • 在四个基准上与人工评分相关性更高,优于现有方法。
  • 适合需要高精度自动评估的文本到音频生成研究者使用。

文本到音频(TTA)生成能够根据自然语言精准捕捉用户意图,但其模型评估依赖昂贵的人工主观评分。为推动发展,需建立与人类判断高度相关的自动评估指标。尽管现有基于CLAP的指标提供无参考解决方案,但其粗粒度的文本-音频匹配常与人工评分相关性不佳。为此,本文提出ELSA,一种无参考的细粒度文本-音频对齐评估方法。ELSA依据文本查询中提取的独立声学事件对生成音频进行分解,并评估事件级别的对齐程度。在四个TTA基准上的实验表明,ELSA相较于先前指标展现出更高的与人工主观评分的相关性,验证了其在可靠评估TTA模型方面的有效性。

原文摘要 · Abstract (English)

Text-to-audio (TTA) generation, synthesizing audio from natural language, has been widely studied for its ability to capture precise user intent. To effectively advance TTA models, it is essential to reliably evaluate generated audio without relying on costly human subjective ratings, motivating the development of automatic evaluation metrics that correlate well with human judgments. While recent CLAP-based metrics provide practical reference-free solutions, their coarse-grained text-audio similarity matching often correlates poorly with human ratings. To address this, we propose ELSA, a reference-free evaluation metric for fine-grained text-audio alignment. ELSA decomposes generated audio guided by distinct acoustic events derived from the text query and assesses event-level alignment. Experiments across four TTA benchmarks show that ELSA reveals a higher correlation with human subjective ratings than prior metrics, highlighting its effectiveness for reliable TTA evaluation.

文本到音频自动评估声学事件细粒度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。