揭示LLM策略研究标注的波动根源,提出可信赖测量协议
Variance-Aware LLM Annotation for Strategy Research: Sources, Diagnostics, and a Protocol for Reliable Measurement
- 从五个维度诊断标注波动:定义、界面、模型偏好等
- 微小设计差异导致结果偏差12%-85%,影响统计推断可靠性
- 提供采样、聚合与报告规范,明确不适用场景
大型语言模型(LLMs)为策略研究提供了大规模文本标注的强大工具,但将模型生成标签视为确定性结果会忽略显著的不稳定性。基于内容分析与普遍性理论,我们诊断出五类方差来源:概念界定、界面效应、模型偏好、输出提取及系统级聚合。实证表明,微小的设计选择——如提示词措辞、模型选型——可导致结果变化12至85个百分点。此类波动不仅威胁可重复性,更影响计量识别:与协变量相关的标注误差会扭曲参数估计,即便平均准确率较高亦然。本文构建了包含采样预算、聚合规则与报告标准的方差感知协议,并明确指出使用LLM标注的适用边界。这些贡献使基于LLM的标注从随意实践转变为可审计的测量基础设施。
原文摘要 · Abstract (English)
Large language models (LLMs) offer strategy researchers powerful tools for annotating text at scale, but treating LLM-generated labels as deterministic overlooks substantial instability. Grounded in content analysis and generalizability theory, we diagnose five variance sources: construct specification, interface effects, model preferences, output extraction, and system-level aggregation. Empirical demonstrations show that minor design choices-prompt phrasing, model selection-can shift outcomes by 12-85 percentage points. Such variance threatens not only reproducibility but econometric identification: annotation errors correlated with covariates bias parameter estimates regardless of average accuracy. We develop a variance-aware protocol specifying sampling budgets, aggregation rules, and reporting standards, and delineate scope conditions where LLM annotation should not be used. These contributions transform LLM-based annotation from ad hoc practice into auditable measurement infrastructure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。