用黄金问题激励标注员产出高质量数据,提升模型训练效果。
Incentivizing High-Quality Human Annotations with Golden Questions
- 设计黄金问题监控标注质量,通过统计检验奖励高质标注。
- 理论证明检验成功率随样本数提升速率为Θ(1/√(n log n))。
- 黄金问题需高确定性且格式贴近日常任务,适合标注质量评估。
人工标注数据在训练大语言模型(如监督微调和人类偏好对齐)中至关重要,但付费标注员未必保证高质量。本文基于委托-代理模型,研究公司(委托方)如何激励标注员(代理方)产出优质数据。委托方仅能通过检查n个样本的标注质量来监控。我们采用最大似然估计(MLE)及相应假设检验机制:若MLE通过检验,代理方获得奖金。分析表明,由于代理的战略行为,该模型的检验速率并非传统的大偏差理论所预测的指数级,而是Θ(1/√(n log n))。理论推导出黄金问题的两个关键标准:(1)高确定性;(2)与常规任务格式一致。据此,我们在人类偏好数据中选取一组黄金问题。通过激励相容实验发现,相比传统的指令操控检测等调查手段,黄金问题更能有效揭示标注员的真实行为。
原文摘要 · Abstract (English)
Human-annotated data plays a vital role in training large language models (LLMs), such as supervised fine-tuning and human preference alignment. However, it is not guaranteed that paid human annotators produce high-quality data. In this paper, we study how to incentivize human annotators to do so. We start from a principal-agent model to model the dynamics between the company (the principal) and the annotator (the agent), where the principal can only monitor the annotation quality by examining $n$ samples. We investigate the maximum likelihood estimators (MLE) and the corresponding hypothesis testing to incentivize annotators: the agent is given a bonus if the MLE passes the test. By analyzing the variance of the outcome, we show that the strategic behavior of the agent makes the hypothesis testing very different from traditional ones: Unlike the exponential rate proved by the large deviation theory, the principal-agent model's hypothesis testing rate is of $Θ(1/\sqrt{n \log n})$. Our theory implies two criteria for the \emph{golden questions} to monitor the performance of the annotators: they should be of (1) high certainty and (2) similar format to normal ones. In that light, we select a set of golden questions in human preference data. By doing incentive-compatible experiments, we find out that the annotators' behavior is better revealed by those golden questions, compared to traditional survey techniques such as instructed manipulation checks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。