用提示词稳定性评估大模型社会科学研究标注的可靠性
What Is Actually Being Annotated? Inter-Prompt Reliability as a Measurement Problem in LLM-Based Social Science Labeling
- 提出跨提示词一致性框架,量化不同表述下模型输出的波动
- 在解释类任务中模型结果随机性显著,知识类任务更稳定
- 多提示投票可大幅提升标注可复现性,适合严谨研究者
大语言模型在计算社会科学标注中应用日益广泛,但其在提示词变化下的方法学可靠性仍不明确。本文提出跨提示词一致性(IPR)框架,通过成对一致率(PAR)及其分布评估模型输出在语义等价但语言形式不同的提示下的稳定性。在两个性质不同的任务上验证:TREC(解释性)和Politifact(知识锚定型)。结果表明,解释性任务中模型标注存在显著随机波动,而知识型任务表现更稳定。进一步发现,对多个提示进行多数投票能显著提升可复现性并降低方差。研究揭示,提示词本身构成一种测量工具,其措辞引入方法学不确定性。建议未来研究摒弃单一提示评估,转向基于IPR框架的分布稳定性分析与提示聚合。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used for annotation in computational social science, yet their methodological reliability under prompt variation remains unclear. This paper introduces Inter-Prompt Reliability (IPR), a framework for evaluating the stability of LLM outputs across semantically equivalent but linguistically varied prompts. Drawing on Inter-Rater Reliability, IPR is measured by Pairwise Agreement Rate (PAR) and its distribution to capture both consistency and stochasticity in model behavior. We evaluate this framework on two tasks with distinct properties: TREC (interpretative) and Politifact (knowledge-anchored). Results show that LLM annotation exhibits substantial stochastic variation in interpretative tasks, while appearing more stable in knowledge-based tasks. We further show that majority voting across prompts significantly improves reproducibility and reduces variance. These findings suggest that LLM prompt acts as an instrumental measurement while its wording exhibits methodological uncertainty. For future LLM-based CSS studies, we suggest that researchers move beyond single-prompt evaluation toward distributional stability and prompt aggregation within our IPR framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。