arXiv:2602.00604cs.SDeess.AS2026-02中稿 · ICASSP 2026 Worksh…

用CLAP伪标签预训练,提升音频文本对齐效果

The TMU System for the XACLE Challenge: Training Large Audio Language Models with CLAP Pseudo-Labels

  • 三阶段训练:自动语音描述预训练+CLAP伪标签预训练+微调
  • 在XACLE测试集上达到0.632的SRCC,远超基线0.334
  • 适合做音频-文本对齐、大模型训练的研究者参考

本文提出针对x-to-audio对齐(XACLE)挑战的参赛方案。目标是预测给定通用音频与文本对之间的语义对齐程度。所提系统基于大型音频语言模型(LALM)架构,采用三阶段训练流程:自动化音频描述预训练、使用CLAP伪标签的预训练,以及在XACLE数据集上的微调。实验表明,使用CLAP伪标签的预训练是性能提升的主要驱动力。在XACLE测试集上,本系统取得0.632的SRCC,显著优于基线系统(0.334),并在挑战赛团队排名中位列第三。代码与模型详见https://github.com/shiotalab-tmu/tmu-xacle2026。

原文摘要 · Abstract (English)

In this paper, we propose a submission to the x-to-audio alignment (XACLE) challenge. The goal is to predict semantic alignment of a given general audio and text pair. The proposed system is based on a large audio language model (LALM) architecture. We employ a three-stage training pipeline: automated audio captioning pretraining, pretraining with CLAP pseudo-labels, and fine-tuning on the XACLE dataset. Our experiments show that pretraining with CLAP pseudo-labels is the primary performance driver. On the XACLE test set, our system reaches an SRCC of 0.632, significantly outperforming the baseline system (0.334) and securing third place in the challenge team ranking. Code and models can be found at https://github.com/shiotalab-tmu/tmu-xacle2026

音频对齐大模型训练伪标签CLAP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。