用轻量人工干预提升大模型标注可靠性,省力45%仍保准确
Reliable Annotations with Less Effort: Evaluating LLM-Human Collaboration in Search Clarifications
- 通过置信度阈值和模型分歧筛选需人工审核的标注
- 相比纯人工,减少45%工作量,且标注一致性显著提升
- 适合需要高精度、低成本标注的复杂评估任务
尽管大语言模型在自动化标注方面备受关注,但在复杂、细微且多维度的标注任务中其效果仍不明确。本研究聚焦搜索澄清任务,利用包含五个细粒度子任务的高质量多维数据集进行评估。结果显示,即使最先进的大模型在主观或细粒度评价任务中也难以达到人类水平表现,其预测常存在不一致、校准不足,并对提示词高度敏感。为此,我们提出一种简单有效的「人机协同」流程:基于置信度阈值与多模型分歧,仅对不确定样本引入人工审查。实验表明,该方法显著提升标注可靠性,同时将人工工作量降低最多45%,为实际评估场景中部署大模型提供了可扩展、低成本且高精度的可行路径。
原文摘要 · Abstract (English)
Despite growing interest in using large language models (LLMs) to automate annotation, their effectiveness in complex, nuanced, and multi-dimensional labelling tasks remains relatively underexplored. This study focuses on annotation for the search clarification task, leveraging a high-quality, multi-dimensional dataset that includes five distinct fine-grained annotation subtasks. Although LLMs have shown impressive capabilities in general settings, our study reveals that even state-of-the-art models struggle to replicate human-level performance in subjective or fine-grained evaluation tasks. Through a systematic assessment, we demonstrate that LLM predictions are often inconsistent, poorly calibrated, and highly sensitive to prompt variations. To address these limitations, we propose a simple yet effective human-in-the-loop (HITL) workflow that uses confidence thresholds and inter-model disagreement to selectively involve human review. Our findings show that this lightweight intervention significantly improves annotation reliability while reducing human effort by up to 45%, offering a relatively scalable and cost-effective yet accurate path forward for deploying LLMs in real-world evaluation settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。