通过融合多模态专家与时间动态建模,提升歌曲流行度预测准确率。
Who Will Top the Charts? Multimodal Music Popularity Prediction via Adaptive Fusion of Modality Experts and Temporal Engagement Modeling
- 设计自适应门控机制融合音频、歌词与社交数据专家
- 在11.3万首歌曲上实现R²达0.69,较基线提升12%
- 适合音乐产业战略决策与内容创作的快速评估
预测歌曲发行前的商业成功仍是音乐产业的关键挑战。早期预测可支持战略决策与营销规划。现有方法存在四大局限:(i) 忽略音频与歌词的时间动态;(ii) 用词袋模型表示歌词,忽略结构与情感语义;(iii) 忽视艺术家与歌曲的历史表现;(iv) 多模态融合依赖简单拼接,导致表征对齐差。为此,本文提出GAMENet,一种端到端多模态深度学习架构。该模型通过自适应门控机制整合音频、歌词与社交元数据的模态专家。音频特征来自Music4AllOnion,经OnionEnsembleAENet(一组用于鲁棒特征提取的自编码器)处理;歌词嵌入通过大语言模型管道生成;引入新特征Career Trajectory Dynamics(CTD),捕捉艺人多年职业动量与歌曲级轨迹统计。基于包含11.3万首歌曲的Music4All数据集(此前用于MIR任务但未用于流行度预测),GAMENet相较直接拼接多模态特征提升12% R²。仅使用Spotify音频描述符时R²为0.13,加入聚合CTD特征后升至0.69,再结合时间维度CTD特征额外提升7%。在10万首歌曲的SpotGenTrack Popularity Dataset上验证,相较之前基线提升16%。大量消融实验确认模型有效性及各模态独立贡献。
原文摘要 · Abstract (English)
Predicting a song's commercial success prior to its release remains an open and critical research challenge for the music industry. Early prediction of music popularity informs strategic decisions, creative planning, and marketing. Existing methods suffer from four limitations:(i) temporal dynamics in audio and lyrics are averaged away; (ii) lyrics are represented as a bag of words, disregarding compositional structure and affective semantics; (iii) artist- and song-level historical performance is ignored; and (iv) multimodal fusion approaches rely on simple feature concatenation, resulting in poorly aligned shared representations. To address these limitations, we introduce GAMENet, an end-to-end multimodal deep learning architecture for music popularity prediction. GAMENet integrates modality-specific experts for audio, lyrics, and social metadata through an adaptive gating mechanism. We use audio features from Music4AllOnion processed via OnionEnsembleAENet, a network of autoencoders designed for robust feature extraction; lyric embeddings derived through a large language model pipeline; and newly introduced Career Trajectory Dynamics (CTD) features that capture multi-year artist career momentum and song-level trajectory statistics. Using the Music4All dataset (113k tracks), previously explored in MIR tasks but not popularity prediction, GAMENet achieves a 12% improvement in R^2 over direct multimodal feature concatenation. Spotify audio descriptors alone yield an R^2 of 0.13. Integrating aggregate CTD features increases this to 0.69, with an additional 7% gain from temporal CTD features. We further validate robustness using the SpotGenTrack Popularity Dataset (100k tracks), achieving a 16% improvement over the previous baseline. Extensive ablations confirm the model's effectiveness and the distinct contribution of each modality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。