arXiv:2509.03116cs.CL2025-09EMNLP被引 16

用大模型测量社会科学研究中的连续语义特征,效果优于直接打分。

Measuring Scalar Constructs in Social Science with LLMs

  • 通过成对比较让大模型评估语义连续性,避免数值聚集问题。
  • 基于词元概率加权平均,测量精度进一步提升。
  • 仅需1000组数据微调小模型,性能可媲美大模型提示法。

许多描述语言的构念(如复杂性、情感性)具有天然的连续语义结构,例如一篇公开演讲并非简单或复杂二分,而处于两者之间的连续谱上。尽管大语言模型(LLMs)是测量此类连续构念的有力工具,但其对数值输出的特定处理方式引发如何有效应用的疑问。我们通过多数据集评估了四种基于LLM的测量方法:无权重的直接点对点评分、成对比较聚合、基于词元概率加权的点对点评分以及微调。研究发现,由LLM进行的成对比较比直接提示输出分数更优,后者常出现数值聚集在任意数字上的问题。进一步地,对词元概率加权后的均值计算显著提升了测量效果。最终,使用仅1,000个训练样本微调小型模型,即可达到甚至超越提示式大模型的表现。

原文摘要 · Abstract (English)

Many constructs that characterize language, like its complexity or emotionality, have a naturally continuous semantic structure; a public speech is not just "simple" or "complex," but exists on a continuum between extremes. Although large language models (LLMs) are an attractive tool for measuring scalar constructs, their idiosyncratic treatment of numerical outputs raises questions of how to best apply them. We address these questions with a comprehensive evaluation of LLM-based approaches to scalar construct measurement in social science. Using multiple datasets sourced from the political science literature, we evaluate four approaches: unweighted direct pointwise scoring, aggregation of pairwise comparisons, token-probability-weighted pointwise scoring, and finetuning. Our study finds that pairwise comparisons made by LLMs produce better measurements than simply prompting the LLM to directly output the scores, which suffers from bunching around arbitrary numbers. However, taking the weighted mean over the token probability of scores further improves the measurements over the two previous approaches. Finally, finetuning smaller models with as few as 1,000 training pairs can match or exceed the performance of prompted LLMs.

大模型社会科学研究连续测量文本分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。