用语言模型分析近70年公告牌歌曲,发现1990年后歌词暴力色情内容显著上升
Language models for longitudinal analysis of abusive content in Billboard Music Charts
- 用深度学习分析美国公告牌70年歌曲歌词,追踪内容演变
- 1990年后歌词中粗俗、性暗示内容明显增多,趋势显著
- 适合关注流行文化变迁与青少年影响的研究者
近年来,音乐中暴力及性暗示内容急剧增加,尤其在公告牌音乐榜(Billboard Music Charts)中表现突出。然而,缺乏有效研究验证这一趋势,难以支持政策制定。本文利用深度学习与语言模型,对过去七十年美国公告牌榜单歌曲的歌词进行纵向分析,结合情感分析与滥用内容检测,系统考察内容演变。结果显示,自1990年起,主流音乐中露骨内容显著上升,包含大量亵渎、性暗示及不当语言的歌曲比例持续增长。研究揭示了语言模型捕捉歌词中细微语义变化的能力,反映了社会规范与语言使用随时间的变迁。
原文摘要 · Abstract (English)
There is no doubt that there has been a drastic increase in abusive and sexually explicit content in music, particularly in Billboard Music Charts. However, there is a lack of studies that validate the trend for effective policy development, as such content has harmful behavioural changes in children and youths. In this study, we utilise deep learning methods to analyse songs (lyrics) from Billboard Charts of the United States in the last seven decades. We provide a longitudinal study using deep learning and language models and review the evolution of content using sentiment analysis and abuse detection, including sexually explicit content. Our results show a significant rise in explicit content in popular music from 1990 onwards. Furthermore, we find an increasing prevalence of songs with lyrics containing profane, sexually explicit, and otherwise inappropriate language. The longitudinal analysis of the ability of language models to capture nuanced patterns in lyrical content, reflecting shifts in societal norms and language use over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。