arXiv:2608.08885physics.soc-phcs.CL2026-08

用大模型量化说唱歌词性内容,可复现且不限于性暗示

Towards an LLM-based method for quantifying the sexual content in song lyrics

论文配图:Towards an LLM-based method for quantifying the sexual content in song lyrics
图 1 · 摘自论文原文
  • 基于大语言模型设计多维度歌词主题量化方法
  • 分析1259首雷鬼音乐歌词,揭示性暗示随时间变化趋势
  • 结果可复现,适合内容分析与流行文化研究者

雷鬼音乐是全球最受欢迎的音乐类型之一,其歌词常被认为高度性化。这一观点主要基于定性研究和小规模定量研究。本文有两个目标:首先,提出一种可复现的方法,利用大语言模型对歌曲歌词中的多个独立主题维度进行量化,该方法不局限于性内容;其次,将该方法应用于1259首由12位雷鬼艺术家在2002至2025年间发布的歌曲。分析涵盖数据集特征、艺术家间比较、各维度随时间演变趋势,以及本研究的性暗示评分与Spotify官方“成人标志”的对比。我们公开了数据收集代码、评分提示和歌词语料库,供其他研究者复现或扩展应用。

原文摘要 · Abstract (English)

Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mostly on qualitative studies and on small-scale quantitative ones. This paper has two goals. First, we present a reproducible method that uses a large language model to quantify thematic content in song lyrics along several independent dimensions. The method is not restricted to sexual content. Second, we apply it to a corpus of 1,259 songs by 12 reggaeton artists released between 2002 and 2025. The analysis covers four topics: a dataset characterization, a per-artist comparison, an analysis of how the dimensions change over time, and a comparison between our sexual-explicitness score and Spotify's own explicit flag. We release the data collection code, the scoring prompt, and the corpus, so that other researchers can replicate the approach or apply it to their own lyrics datasets.

大模型歌词分析性内容量化可复现研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。