用图模型发现K-pop歌词中的隐含主题群,揭示跨主题歌曲的独特语言特征。
Semantic Communities and Boundary-Spanning Lyrics in K-pop: A Graph-Based Unsupervised Analysis
- 基于歌词级语义构建相似性图,无监督发现稳定主题社区。
- 跨社区歌曲词汇多样性更高,重复率更低,打破重复驱动连接的假设。
- 方法不依赖语言、流派或歌手标签,适合分析无标注文化文本。
大规模歌词语料库在数据驱动分析中面临标注不可靠、多语言内容和高风格重复等挑战。现有方法多依赖有监督分类、流派标签或粗粒度文档表示,难以揭示潜在语义结构。本文提出一种基于图的无监督框架,通过构建歌词文本间的相似性图并应用社区检测,发现无需流派、艺术家或语言监督的稳定微主题社区。进一步利用图论桥接度量识别跨边界歌曲,并分析其结构特性。在多种鲁棒性设置下,跨社区歌词表现出更高的词汇多样性与更低的重复率,挑战了‘副歌强度或重复性驱动跨主题连通性’的假设。该框架具有语言无关性,适用于未标注的文化文本语料库。
原文摘要 · Abstract (English)
Large-scale lyric corpora present unique challenges for data-driven analysis, including the absence of reliable annotations, multilingual content, and high levels of stylistic repetition. Most existing approaches rely on supervised classification, genre labels, or coarse document-level representations, limiting their ability to uncover latent semantic structure. We present a graph-based framework for unsupervised discovery and evaluation of semantic communities in K-pop lyrics using line-level semantic representations. By constructing a similarity graph over lyric texts and applying community detection, we uncover stable micro-theme communities without genre, artist, or language supervision. We further identify boundary-spanning songs via graph-theoretic bridge metrics and analyse their structural properties. Across multiple robustness settings, boundary-spanning lyrics exhibit higher lexical diversity and lower repetition compared to core community members, challenging the assumption that hook intensity or repetition drives cross-theme connectivity. Our framework is language-agnostic and applicable to unlabeled cultural text corpora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。