arXiv:2504.03612cs.CL2025-04中稿 · COLM被引 5

拆解提示词数据集三要素,提升大模型对齐效率

AIR: A Systematic Analysis of Annotations, Instructions, and Response Pairs in Preference Dataset

  • 分离标注、指令、回复三要素,分别优化
  • 仅用1.4万优质样本,性能提升5.3%以上
  • 适合追求高效对齐的模型训练团队

偏好学习对齐大语言模型与人类价值观至关重要,其成功依赖于标注、指令和回复对三个核心组件。现有方法混同这些成分,掩盖各自影响,阻碍系统优化。本文提出AIR框架,系统性地分离并优化各组件,评估其协同效应。实验表明:标注应简洁(逐项生成评分)、指令需稳定(基于方差跨模型过滤)、回复对质量要求中等间隔加高绝对分值。三者结合使基准方法平均提升5.3%,即便仅使用1.4万高质量样本。本工作将偏好数据集设计从随意扩增转向组件感知优化,为高效可复现对齐提供蓝图。

原文摘要 · Abstract (English)

Preference learning is critical for aligning large language models (LLMs) with human values, yet its success hinges on high-quality datasets comprising three core components: Preference \textbf{A}nnotations, \textbf{I}nstructions, and \textbf{R}esponse Pairs. Current approaches conflate these components, obscuring their individual impacts and hindering systematic optimization. In this work, we propose \textbf{AIR}, a component-wise analysis framework that systematically isolates and optimizes each component while evaluating their synergistic effects. Through rigorous experimentation, AIR reveals actionable principles: annotation simplicity (point-wise generative scoring), instruction inference stability (variance-based filtering across LLMs), and response pair quality (moderate margins + high absolute scores). When combined, these principles yield +5.3 average gains over baseline method, even with only 14k high-quality pairs. Our work shifts preference dataset design from ad hoc scaling to component-aware optimization, offering a blueprint for efficient, reproducible alignment.

偏好学习数据集优化大模型对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。