用五种大模型分析英文歌词中的文化心理,验证其标注可靠性。
A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models

- 通过重复标注检验五种大模型在四种心理特质上的稳定性。
- 自尊标注最可靠,寻求认可的标注最不稳定。
- 适合做文化分析的大模型需报告重复性与跨模型一致性。
大型语言模型(LLMs)正被用于大规模标注文化文本,但其输出是否可作为社会隐变量的可靠测量仍需验证。本研究将五种LLM作为零样本标注器,分析英语歌曲歌词中四种社会心理特质:自尊、自我控制、归属感需求和认可需求。基于对大规模歌词语料的重复标注,考察了三种测量特性:多次运行的一致性、跨模型收敛性以及共识标签在监督分类中的可迁移性。结果显示,不同心理特质的标注可靠性差异显著:自尊具有最强的重复测量稳定性,而寻求认可的标注普遍较不稳定;自我控制和归属感需求表现居中,且依赖具体模型。下游分类实验表明,共识标签包含可学习信号,但可迁移性本身不等同于构念效度。因此,在文化分析中使用LLM标注前,必须报告重复测量稳定性和跨模型一致性。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used to annotate cultural texts at scales that are impractical for human coders. However, before their outputs are treated as measurements of latent social constructs, it is necessary to establish whether those measurements are reliable. This study evaluates five LLMs as zero-shot annotators of four social constructs expressed in English song lyrics: self-esteem, self-control, seeking belonging, and seeking recognition. Using repeated annotations of a large lyric corpus, we examine three properties of LLM-based measurement: consistency across repeated runs, convergence across models, and transferability of consensus labels to supervised classification. The findings show that LLM-based measurement is not uniformly reliable across constructs. Self-esteem exhibits the strongest repeated-measurement reliability across models, while seeking recognition is generally less stable; self-control and seeking belonging show intermediate but model-dependent reliability. Downstream classification further indicates that consensus LLM labels contain learnable signal, although transferability does not itself establish construct validity. Repeated-measurement stability and cross-model convergence should therefore be reported before LLM annotations are treated as scalable measurements in cultural analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。