arXiv:2502.06387cs.LGcs.GT2025-02被引 5

提出新方法评估和激励人工标注者,提升大模型对齐质量。

How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators

  • 用自洽性监控替代专家审核,更适配偏好标注特点。
  • 实证表明自洽性监控在样本量较小时效果优于传统方法。
  • 设计合同机制,让标注质量可量化并激励长期高质量产出。

人工标注的偏好数据在大语言模型对齐中至关重要。本文研究两个核心问题:如何监控标注者质量,以及如何激励其提供高质量标注。现有基于专家的监控在偏好标注中表现不佳,因标注者差异大,且下游模型性能是噪声强烈的间接指标。为此,本文提出专为偏好标注设计的自洽性监控方案,并分析两种方法的统计样本复杂度。结果揭示了可靠评估标注者所需的最小样本数,并说明自洽性监控何时更优。进一步,将监控信号用于主代理模型,研究在何种样本量下简单合约可逼近理想表现。在连续动作空间下,二值合约误差率收敛速度为 $Θ(1/\\(sqrt{\mathcal{I} n \log n})$,线性合约为 $Θ(1/(\mathcal{I}n))$,且后者在一般合约中达到最优速率。这与离散动作空间下二值合约呈指数衰减的已知结果形成对比。

原文摘要 · Abstract (English)

Human-annotated preference data play an important role in aligning large language models (LLMs). In this paper, we study two connected questions: how to monitor the quality of human preference annotators and how to incentivize them to provide high-quality annotations. In current practice, expert-based monitoring is a natural workhorse for quality control, but it performs poorly in preference annotation because annotators are heterogeneous and downstream model performance is an indirect and noisy proxy for annotation quality. We therefore propose a self-consistency monitoring scheme tailored to preference annotation, and analyze the statistical sample complexity of both methods. This practitioner-facing analysis identifies how many inspected samples are needed to reliably assess an annotator and shows when self-consistency monitoring can outperform expert-based monitoring. We then use the resulting monitoring signal as the performance measure in a principal-agent model, which lets us study a second sample-complexity question: how many monitored samples are needed before simple contracts perform close to the ideal benchmark in which annotation quality is perfectly observable. Under this continuous action space, we show that this shortfall scales as $Θ(1/\sqrt{\mathcal{I} n \log n})$ for binary contracts and $Θ(1/(\mathcal{I}n))$ for linear contracts, where $\mathcal{I}$ is the Fisher information and $n$ is the number of samples; we further show that the linear contracts are rate-optimal among general contracts. This contrasts with the known result that binary contracts are optimal and of $\exp(-Θ(n))$ when the action space is discrete \citep{frick2023monitoring}.

大模型对齐标注质量激励机制自洽性监控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。