为学术界打造公平的文本生成音乐挑战,推动开源模型发展
Academic Text-to-Music Grand Challenge: Datasets, Baselines, and Evaluation Methods
- 用开源音乐数据集从零训练模型,限制参数量保证公平性
- 采用FAD、CLAP和概念覆盖度等多指标评估,结合听觉测试
- 适合关注文本到音乐生成的学者与开源研究者
本文介绍了ICME 2026学术文本到音乐生成(ATTM)大挑战的总体框架。尽管文本到音乐生成(TTM)系统进展迅速,但当前领域主要由基于海量专有数据和工业级算力训练的模型主导,阻碍了学术研究。为此,ATTM挑战建立了一个公平基准,要求参赛者仅使用标准化、CC授权的MTG-Jamendo数据子集(仅含器乐)从头训练生成模型。挑战分为效率赛道(限定500M参数)和性能赛道(无参数限制)。提交结果通过多阶段评估:先用客观指标(包括弗雷切特音频距离、CLAP得分和新提出的概念覆盖度评分,即CCS),再进行主观听觉测试。本挑战提供开源基线、预处理流程、参考标题及计算FAD和CLAP的公开代码,旨在促进学术环境下的TTM研究。
原文摘要 · Abstract (English)
This paper presents an overview and the technical framework of the ICME 2026 Grand Challenge on Academic Text-to-Music Generation (ATTM). Despite the rapid progress in text-to-music generation (TTM) systems, the field is currently dominated by models trained on massive proprietary datasets with industrial-scale computational resources, creating a significant barrier for academic research. To address this, the ATTM Challenge establishes a fair-play benchmark that requires participants to train generative models strictly from scratch using a standardized, CC-licensed subset of the MTG-Jamendo dataset containing only instrumental music. The challenge is divided into two tracks: the Efficiency Track (limited to 500M parameters) and the Performance Track (no parameter limit). Submissions are evaluated through a multi-stage process involving objective metrics, including Frechet Audio Distance, CLAP score, and a novel Concept Coverage Score (CCS), followed by a subjective listening test. By providing open-source baselines, preprocessing pipelines, reference captions, and public evaluation code for computing FAD and CLAP, this challenge aims to facilitate and promote TTM research in academic contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。