arXiv:2601.06829cs.SD2026-01

用专家混合模型评估文本到音频生成的语义一致性,准确率超基线30.6%。

MoEScore: Mixture-of-Experts-Based Text-Audio Relevance Score Prediction for Text-to-Audio System Evaluation

  • 采用专家混合架构结合序列交叉注意力,自动判断音频与文本匹配度。
  • 在XACLE挑战赛测试集上相关系数达0.6402,领先基线30.6%。
  • 适合需要客观评估语音生成质量的研究者和开发者使用。

生成模型的进步使现代文本到音频(TTA)系统能够合成高感知质量的音频。然而,TTA系统常难以保持输入文本的语义一致性,导致声音事件、时间结构或上下文关系不匹配。评估TTA中的语义保真度仍是重大挑战。传统方法主要依赖耗时的人工听觉测试。为此,我们提出一种基于专家混合(MoE)架构与序列交叉注意力(SeqCoAttn)的客观评估模型。该模型在XACLE挑战赛中排名第一,在测试集上达到SRCC 0.6402,较挑战基线提升30.6%。代码已公开于:https://github.com/S-Orion/MOESCORE。

原文摘要 · Abstract (English)

Recent advances in generative models have enabled modern Text-to-Audio (TTA) systems to synthesize audio with high perceptual quality. However, TTA systems often struggle to maintain semantic consistency with the input text, leading to mismatches in sound events, temporal tructures, or contextual relationships. Evaluating semantic fidelity in TTA remains a significant challenge. Traditional methods primarily rely on subjective human listening tests, which is time-consuming. To solve this, we propose an objective evaluator based on a Mixture of Experts (MoE) architecture with Sequential Cross-Attention (SeqCoAttn). Our model achieves the first rank in the XACLE Challenge, with an SRCC of 0.6402 (an improvement of 30.6% over the challenge baseline) on the test dataset. Code is available at: https://github.com/S-Orion/MOESCORE.

音频生成语义评估专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。