OMAR-RQ用33万小时音乐数据训练,实现开源音乐表征新标杆。
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
- 基于多特征掩码预测自监督训练,构建通用音乐表征模型。
- 在音乐标签、和弦识别等6项任务上达开源模型最佳性能。
- 适合音乐信息检索与跨任务迁移研究者使用。
构建开源基础模型对于推进音乐音频理解研究至关重要,能为音乐信息检索提供强大且通用的表示能力。我们提出OMAR-RQ,该模型通过自监督学习方法,利用超过33万小时的音乐音频大规模数据集进行训练,采用掩码标记分类策略。我们测试了多种输入特征与量化方案,在音乐标签、音高估计、和弦识别、节拍追踪、分割和难度评估等多项任务中,性能优于现有开源自监督模型。我们已开源训练与评估流程及模型权重,详见https://github.com/mtg/omar-rq。
原文摘要 · Abstract (English)
Developing open-source foundation models is essential for advancing research in music audio understanding and ensuring access to powerful, multipurpose representations for music information retrieval. We present OMAR-RQ, a model trained with self-supervision via masked token classification methodologies using a large-scale dataset with over 330,000 hours of music audio. We experiment with different input features and quantization options, and achieve state-of-the-art performance in music tagging, pitch estimation, chord recognition, beat tracking, segmentation, and difficulty estimation among open self-supervised models. We open-source our training and evaluation pipelines and model weights, available at https://github.com/mtg/omar-rq.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。