arXiv:2601.16273cs.SDeess.AS2026-01被引 1

基于BEATs的音频编码器,用74000小时数据训练,性能超基线和Dasheng模型。

The CMU-AIST submission for the ICME 2025 Audio Encoder Challenge

  • 用7.4万小时多源数据扩展BEATs,模型达3亿参数。
  • 在语音与音效混合数据上训练,提升音频表征能力。
  • 提出简单集成方法,显著优于基线和12亿参数模型。

本文描述了我们对ICME 2025音频编码器挑战赛的提交方案。系统基于基于掩码语音标记预测的音频编码器BEATs,利用来自多种语音、音乐和声音语料库的74,000小时数据进行扩展,并将模型架构扩展至3亿参数。我们实验了以语音为主和均衡混合的预训练数据组合,研究不同领域对最终性能的影响。提交系统由一个12亿参数的Dasheng模型与两个在上述预训练数据组合上训练的定制化放大版BEATs模型组成的集成模型构成。我们还提出一种简单的集成技术,有效保留各子模型优势,超越基线及Dasheng 1.2B模型。为促进开放科学,我们在Hugging Face公开发布了训练好的检查点:https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND 和 https://huggingface.co/shikhar7ssu/OpenBEATs-ICME。

原文摘要 · Abstract (English)

This technical report describes our submission to the ICME 2025 audio encoder challenge. Our submitted system is built on BEATs, a masked speech token prediction based audio encoder. We extend the BEATs model using 74,000 hours of data derived from various speech, music, and sound corpora and scale its architecture upto 300 million parameters. We experiment with speech-heavy and balanced pre-training mixtures to study the impact of different domains on final performance. Our submitted system consists of an ensemble of the Dasheng 1.2 billion model with two custom scaled-up BEATs models trained on the aforementioned pre-training data mixtures. We also propose a simple ensembling technique that retains the best capabilities of constituent models and surpasses both the baseline and Dasheng 1.2B. For open science, we publicly release our trained checkpoints via huggingface at https://huggingface.co/shikhar7ssu/OpenBEATs-ICME-SOUND and https://huggingface.co/shikhar7ssu/OpenBEATs-ICME.

音频编码BEATs模型集成开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。