用大模型内部状态生成可复现的社会科学文本度量方法
LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research
- 通过大模型隐藏状态学习概念向量,投影计算文本概念值
- 三组复现研究验证了度量的高有效性和可重复性
- 适合需要标准化文本分析的社会科学研究者
社会科学研究中越来越多地使用文本作为数据,亟需开发出有效、一致、可复现且高效的文本概念度量方法。本文提出一种新方法,利用大语言模型(LLM)的内部隐藏状态生成这些概念度量。具体而言,该方法学习一个能捕捉大模型内部目标概念表征的概念向量,并通过将文本的LLM隐藏状态投影到该概念向量上来估计其概念值。三个复现研究证明了该方法在多种社会科学研究情境下均能生成高度有效、一致且可复现的文本度量,显示出其作为研究工具的巨大潜力。
原文摘要 · Abstract (English)
The increasing use of text as data in social science research necessitates the development of valid, consistent, reproducible, and efficient methods for generating text-based concept measures. This paper presents a novel method that leverages the internal hidden states of large language models (LLMs) to generate these concept measures. Specifically, the proposed method learns a concept vector that captures how the LLM internally represents the target concept, then estimates the concept value for text data by projecting the text's LLM hidden states onto the concept vector. Three replication studies demonstrate the method's effectiveness in producing highly valid, consistent, and reproducible text-based measures across various social science research contexts, highlighting its potential as a valuable tool for the research community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。