首个针对AI生成音乐的多任务流行度预测框架,结合审美质量提升预测准确率。
APEX: Large-scale Multi-task Aesthetic-Informed Popularity Prediction for AI-Generated Music

- 基于211k首歌曲构建多任务模型,联合预测播放量、点赞数与五维审美特征。
- 在未见过的生成音乐系统上测试,加入审美特征使偏好预测准确率显著提升。
- 适合音乐推荐、平台内容评估及生成模型优化的研究者与开发者参考。
音乐流行度预测日益受到关注,对艺术家、平台和推荐系统均有意义。然而,AI生成音乐平台的爆发式增长带来了全新且尚未充分探索的场景:每天大量歌曲被生成与消费,却缺乏传统意义上的艺人声誉或唱片公司背书。在此背景下,审美质量这一关键因素仍鲜有研究。我们提出APEX,首个大规模多任务学习框架,用于AI生成音乐的流行度预测。该框架在来自Suno和Udio的超过211,000首歌曲(约10,000小时音频)上训练,联合预测基于互动的流行度信号——播放量与点赞评分——同时预测五个从冻结的MERT音频嵌入中提取的感知审美维度。审美质量与流行度捕捉音乐的不同方面,二者协同具有价值:在包含11个未见生成音乐系统的Music Arena数据集上进行分布外评估,结果显示,加入审美特征能持续提升偏好预测性能,证明所学表征在不同生成架构间具备强泛化能力。
原文摘要 · Abstract (English)
Music popularity prediction has attracted growing research interest, with relevance to artists, platforms, and recommendation systems. However, the explosive rise of AI-generated music platforms has created an entirely new and largely unexplored landscape, where a surge of songs is produced and consumed daily without the traditional markers of artist reputation or label backing. Key, yet unexplored in this pursuit is aesthetic quality. We propose APEX, the first large-scale multi-task learning framework for AI-generated music, trained on over 211k songs (10k hours of audio) from Suno and Udio, that jointly predicts engagement-based popularity signals - streams and likes scores - alongside five perceptual aesthetic quality dimensions from frozen audio embeddings extracted from MERT, a self-supervised music understanding model. Aesthetic quality and popularity capture complementary aspects of music that together prove valuable: in an out-of-distribution evaluation on the Music Arena dataset, comprising pairwise human preference battles across eleven generative music systems unseen during training, including aesthetic features consistently improves preference prediction, demonstrating strong generalisation of the learned representations across generative architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。