arXiv:2512.15979cs.SEcs.AI2025-12被引 3

提出OLAF框架,让大模型标注更可靠可复现。

OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering

  • 将大模型标注视为测量过程,建立可靠性等五大核心概念
  • 强调标注需标准化,避免结果不可靠或无法复现
  • 适合关注大模型在软件工程中应用的科研人员

大语言模型(LLMs)在经验软件工程(ESE)中被越来越多地用于自动化或辅助标注任务,如提交记录、问题和定性数据的标记。然而,这类标注的可靠性与可复现性仍缺乏深入研究。现有研究常缺少对可靠性、校准和漂移的标准度量,且常遗漏关键配置细节。本文认为,基于LLM的标注应被视为一种测量过程,而非单纯的自动化操作。我们提出了一个概念性框架——面向大模型标注的可操作化框架(OLAF),整合了可靠性、校准、漂移、共识、聚合和透明性等关键构建。本文旨在推动方法论讨论,并促进未来在软件工程研究中实现更透明、可复现的基于大模型的标注工作。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used in empirical software engineering (ESE) to automate or assist annotation tasks such as labeling commits, issues, and qualitative artifacts. Yet the reliability and reproducibility of such annotations remain underexplored. Existing studies often lack standardized measures for reliability, calibration, and drift, and frequently omit essential configuration details. We argue that LLM-based annotation should be treated as a measurement process rather than a purely automated activity. In this position paper, we outline the \textbf{Operationalization for LLM-based Annotation Framework (OLAF)}, a conceptual framework that organizes key constructs: \textit{reliability, calibration, drift, consensus, aggregation}, and \textit{transparency}. The paper aims to motivate methodological discussion and future empirical work toward more transparent and reproducible LLM-based annotation in software engineering research.

大模型标注软件工程可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。