arXiv:2512.05592cs.SDeess.AS2025-12中稿 · IEEE ASRU 2025

用KAN和VERSA模型构建音频美学评分系统,效果优于所有提交方案。

The T12 System for AudioMOS Challenge 2025: Audio Aesthetics Score Prediction System Using KAN- and VERSA-based Models

  • 用分组有理KAN替代传统网络层,结合真实与伪标签训练
  • 在语句级三项、系统级两项及整体平均上相关性领先
  • 适合关注音频质量评估与模型创新的研究者

我们为音频美学评分挑战赛2025(AMC25)第2赛道提出由CyberAgent开发的音频美学评分(AES)预测系统AESCA。该系统包含基于科尔莫戈罗夫-阿诺德网络(KAN)的音频盒子美学预测器和基于VERSA工具包的评分预测器。在KAN预测器中,将基线模型中的每层多层感知机替换为分组有理KAN,并使用标注与伪标注音频样本进行训练。VERSA预测器采用极端梯度提升回归模型,融合现有度量输出。两个模型均预测包括四项评估维度在内的音频美学得分。最终通过集成四个KAN模型与一个VERSA模型生成综合评分。所提T12系统在语句级三项、系统级两项以及整体平均的相关性上均优于所有提交系统。同时,我们公开了所提KAN预测器(KAN #1-#4)的推理模型。

原文摘要 · Abstract (English)

We propose an audio aesthetics score (AES) prediction system by CyberAgent (AESCA) for AudioMOS Challenge 2025 (AMC25) Track 2. The AESCA comprises a Kolmogorov--Arnold Network (KAN)-based audiobox aesthetics and a predictor from the metric scores using the VERSA toolkit. In the KAN-based predictor, we replaced each multi-layer perceptron layer in the baseline model with a group-rational KAN and trained the model with labeled and pseudo-labeled audio samples. The VERSA-based predictor was designed as a regression model using extreme gradient boosting, incorporating outputs from existing metrics. Both the KAN- and VERSA-based models predicted the AES, including the four evaluation axes. The final AES values were calculated using an ensemble model that combined four KAN-based models and a VERSA-based model. Our proposed T12 system yielded the best correlations among the submitted systems, in three axes at the utterance level, two axes at the system level, and the overall average. We also released the inference model of the proposed KAN-based predictor (KAN #1-#4).

音频评估KANVERSA评分系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。