LLM测社会现象时信心常不准,该方法让模型更靠谱。
Assessing and Mitigating Miscalibration in LLM-Based Social Science Measurement
- 用大模型输出+信心值生成软标签,训练小分类器校准
- 平均降低43.2%误差(ECE),34.0%布里尔分数
- 适合做文本测量的社科研究者,尤其关注结果可信度
大型语言模型(LLMs)正被广泛用于社会科学领域,作为将非结构化文本转化为可进入标准实证分析变量的高效工具。但测量有效性不仅依赖高平均准确率,更需模型置信度真实反映其判断正确的概率。本文研究了基于LLM的社会科学测量中的模型置信度失准问题。以联邦公开市场委员会(FOMC)为例,发现基于信心的筛选会因模型置信度偏差而改变下游回归结果。随后,我们在14个社会科学研究主题上审计了包括GPT-5-mini、DeepSeek-V3.2在内的专有模型与开源模型的校准情况,发现各任务与模型族中,报告置信度与基于容忍度的正确率严重不匹配。为此提出一种软标签蒸馏流程:将大模型的得分及其语义化信心转换为软目标分布,并用该目标训练较小的判别性分类器。在多个数据集上平均使经验校准误差(ECE)降低43.2%,布里尔分数(Brier)降低34.0%。结果表明,校准应作为社会科学研究中基于LLM测量的有效性核心部分,而非可选后处理步骤。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in social science as scalable measurement tools for converting unstructured text into variables that can enter standard empirical designs. Measurement validity demands more than high average accuracy, which requires well calibrated confidence that faithfully reflects the empirical probability of each measurement being correct. This paper studies the model miscalibration in LLM-based social science measurement. We begin with a case study on FOMC and show that confidence based filtering can change downstream regression estimates when LLM confidence is miscalibrated. We then audit calibration across 14 social science constructs covering both proprietary models, including GPT-5-mini, DeepSeek-V3.2, and open source models. Across tasks and model families, reported confidence is poorly aligned with tolerance-based correctness. As a simple mitigation, we propose a soft label distillation pipeline for calibrating Bert with LLM. The method converts an LLM score and its verbalized confidence into a soft target distribution, then trains a smaller discriminative classifier on encoder models for these targets. Averaged across datasets, this approach reduces ECE by 43.2\% and Brier by 34.0\%. These results suggest that LLM-based social science pipelines should treat calibration as part of measurement validity, rather than as an optional post-processing concern.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。