arXiv:2502.18744cs.AIcs.CL2025-02EMNLP

用模型表现知识生成无标注偏好数据,省去人工标注成本。

ZEBRA: Leveraging Model-Behavioral Knowledge for Zero-Annotation Preference Dataset Construction

  • 基于模型在基准测试中的表现判断回复优劣,自动构建偏好对。
  • 在无需任何标注的情况下,对齐效果接近传统人工标注方法。
  • 适合需要大规模低成本对齐数据的研究者或工业应用。

近期大语言模型对齐研究依赖人类或人工智能标注构建大规模偏好数据集,但此类方法需逐样本监督,成本高昂且可解释性差。本文提出ZEBRA——一种基于模型行为的零标注框架,通过分析模型在基准测试中的表现来推断响应质量与相似性,进而二值化生成偏好数据对,完全跳过实例级标注。该方法实现可扩展、可控制、低成本的对齐数据生成。实验证明,尽管无需人工或模型标注,ZEBRA仍能达到与实例监督方法相当的对齐性能。

原文摘要 · Abstract (English)

Recent efforts in LLM alignment have focused on constructing large-scale preference datasets via human or Artificial Intelligence (AI) annotators. However, such approaches rely on instance-wise supervision, incurring substantial annotation cost and limited interpretability. In this paper, we propose ZEBRA - a model behavior-wise zero-annotation framework that constructs preference data by leveraging model behavior knowledge derived from benchmark performances. ZEBRA binarizes response pairs by evaluating the quality and similarity of their origin models, entirely bypassing instance-level annotation. This allows scalable, controllable, and cost-effective alignment data generation. Empirical results show that ZEBRA achieves alignment performance comparable to instance-supervised methods, despite requiring no manual or model-based labeling.

对齐偏好数据零标注模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。