用随机量化压缩音乐理解模型,性能不降反优。
Linear Complexity Self-Supervised Learning for Music Understanding with Random Quantizer
- 结合分支结构与随机量化,降低模型复杂度。
- 模型规模缩小8.5%至12.3%,性能仍达顶尖水平。
- 适合资源受限场景下的音乐信息检索应用。
近年来,基础模型因在自然语言处理任务中表现出色而广受欢迎,通常包含数亿甚至数百亿参数,导致训练和部署成本高昂。本文聚焦于音乐信息检索(MIR)任务中基础模型的规模缩减问题。研究将分支结构(Branchformer)与摘要混合机制(SummaryMixing)结合,并引入随机量化过程。为保证可复现性,我们在公开数据集上进行预训练,并使用一个规模与文献中其他私有数据集相当的专有数据集。通过涵盖多种下游MIR任务的评估框架,验证了模型的鲁棒性。结果表明,该架构在保持与采用多头自注意力的先进模型相当性能的同时,将模型规模减少了8.5%至12.3%。
原文摘要 · Abstract (English)
In recent years, foundation models have become very popular due to their exceptional performance, mainly in natural language (NLP) tasks where they were first introduced. These models usually consist of hundreds of millions, or even billions, of parameters, making them resource-intensive during training and in production systems, leading to increased costs. This paper focuses on the reduction of a foundation's model size when applied to music information retrieval (MIR) tasks. Our research combines the Branchformer architecture with SummaryMixing, which were first applied in speech recognition, along with a random quantization process. To facilitate reproducibility, we conduct pre-training on publicly available datasets, complemented by a proprietary dataset comparable in scale to other private datasets reported in the literature. We ensure robust evaluation by using a framework consisting of a variety of downstream MIR tasks. Our results show that our architecture achieves competitive performance when compared with other state-of-the-art models that use multi-head self-attention, while reducing the model size from 8.5% up to 12.3%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。