arXiv:2412.14461cs.CL2024-12

解决大模型标注的可复现性难题,让标注结果不因模型退役而失效。

To Err Is Human; To Annotate, SILICON? Toward Robust Reproducibility in LLM Annotation

  • 分解标注误差为四类来源,针对性设计干预策略
  • 实证显示误差降低,下游统计估计更准确
  • 提出开源备份模型与智能路由,保障长期可用性

非结构化文本标注是管理研究的基础。大模型虽成本低、可扩展,但其自身可能被弃用,威胁长期可复现性。为此,需实现鲁棒可复现性——当原始模型不可用时仍能复现结果。本文构建分析框架,将测量误差分解为:准则引发误差(标注标准不一)、基线引发误差(人工参考不可靠)、提示引发误差(元指令不佳)和模型引发误差(架构差异)。提出SILICON工作流,在每类误差源上实施针对性干预。在九项管理研究任务中验证,干预显著降低误差,模拟显示其提升下游统计估计精度。进一步提出基于回归的备份开放权重模型方法,所有测试任务均找到性能无显著差异的开源模型。最后提出路由机制,选择性将低置信度样本送至辅助模型,量化当前模型集所能达到的标注质量上限,揭示聚合何时增效、何时反噬。

原文摘要 · Abstract (English)

Unstructured text data annotation is foundational to management research. LLMs offer a cost-effective and scalable alternative to human annotation, but they introduce a novel challenge: the annotator itself can be retired. Proprietary models undergo regular deprecation cycles, threatening long-term reproducibility. Hence, the ability to reproduce annotation results when the original model becomes unavailable, i.e., robust reproducibility, is a central methodological challenge for LLM-based annotation. Achieving robust reproducibility requires first controlling measurement error. We develop an analytical framework that decomposes measurement error into four sources: guideline-induced error from inconsistent annotation criteria, baseline-induced error from unreliable human references, prompt-induced error from suboptimal meta-instruction, and model-induced error from architectural differences across LLMs. We develop the SILICON workflow that instantiates the analytical framework, prescribing targeted interventions at each error source. Empirical validation across nine management research tasks confirms that these interventions reduce measurement error, and simulations show that the resulting error reduction yields more accurate downstream statistical estimates. With measurement error controlled, we address two further aspects of robust reproducibility. First, we propose a regression-based methodology to establish backup open-weight models, which are permanently accessible. Every tested task has at least one open-weight model with no statistically detectable performance difference. Second, we quantify the upper bound of annotation quality attainable from the current set of available models by proposing a routing procedure that selectively sends low-confidence items to auxiliary models, revealing when model aggregation improves performance and when that may adversely affect labeling quality.

大模型标注可复现性误差控制开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。