提出新方法评估大模型在社会科学研究中的分类置信度。
LLM Confidence Evaluation Measures in Zero-Shot CSS Classification
- 设计专用于数据标注的不确定性量化评估指标。
- 对比五种策略,验证新聚合方法能更好识别低置信度标注。
- 适合需要人机协作标注的研究者使用。
在计算社会科学(CSS)任务中,准确评估分类置信度对利用大语言模型(LLMs)进行自动化标注至关重要。本文提出三项关键贡献:(1)设计一种针对数据标注任务的不确定性量化(UQ)性能评估指标;(2)首次在三种不同LLM和三个CSS数据标注任务上比较了五种不同的UQ策略;(3)引入一种新颖的UQ聚合策略,能有效识别低置信度的LLM标注,并显著发现模型误标的数据。实验结果表明,该聚合策略优于现有方法,可显著提升人机协作数据标注流程的效率与准确性。
原文摘要 · Abstract (English)
Assessing classification confidence is critical for leveraging large language models (LLMs) in automated labeling tasks, especially in the sensitive domains presented by Computational Social Science (CSS) tasks. In this paper, we make three key contributions: (1) we propose an uncertainty quantification (UQ) performance measure tailored for data annotation tasks, (2) we compare, for the first time, five different UQ strategies across three distinct LLMs and CSS data annotation tasks, (3) we introduce a novel UQ aggregation strategy that effectively identifies low-confidence LLM annotations and disproportionately uncovers data incorrectly labeled by the LLMs. Our results demonstrate that our proposed UQ aggregation strategy improves upon existing methods andcan be used to significantly improve human-in-the-loop data annotation processes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。