用大模型提取歌词语义,提升流行音乐预测准确率
Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction
- 用大模型生成高维歌词嵌入,捕捉语义与结构信息
- 在超十万首歌曲数据上,误差比基线降低9%至20%
- 适合音乐产业从业者和流媒体平台做内容推荐
精准预测音乐流行度是音乐行业的重要挑战,有助于艺术家、制作人及流媒体平台。以往研究多聚焦音频特征、社交元数据或模型架构,而忽略了歌词的作用。本文提出自动化流程,利用大语言模型(LLM)提取高维歌词嵌入,捕获语义、句法与序列信息。这些特征被集成到多模态模型 HitMusicLyricNet 中,结合音频、歌词与社交元数据,预测流行度得分(0-100)。在包含超过10万首歌曲的 SpotGenTrack 数据集上,该方法相较现有基线在 MAE 上提升9%,在 MSE 上提升20%。消融实验表明,性能提升主要来自由 LLM 驱动的歌词特征管道(LyricsAENet),凸显密集歌词表示的价值。
原文摘要 · Abstract (English)
Accurately predicting music popularity is a critical challenge in the music industry, offering benefits to artists, producers, and streaming platforms. Prior research has largely focused on audio features, social metadata, or model architectures. This work addresses the under-explored role of lyrics in predicting popularity. We present an automated pipeline that uses LLM to extract high-dimensional lyric embeddings, capturing semantic, syntactic, and sequential information. These features are integrated into HitMusicLyricNet, a multimodal architecture that combines audio, lyrics, and social metadata for popularity score prediction in the range 0-100. Our method outperforms existing baselines on the SpotGenTrack dataset, which contains over 100,000 tracks, achieving 9% and 20% improvements in MAE and MSE, respectively. Ablation confirms that gains arise from our LLM-driven lyrics feature pipeline (LyricsAENet), underscoring the value of dense lyric representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。