提出新方法提升大模型对齐的数据效率,减少人工标注成本。
Nearly Optimal Active Preference Learning and Its Application to LLM Alignment
- 基于偏好学习特性设计新算法,突破传统实验设计局限
- 首个实现实例相关标签复杂度保证的主动学习方法
- 实际数据集上显著降低标注需求,适合资源有限的对齐研究
大语言模型对齐依赖高质量的人类偏好数据集,但其收集成本高昂。尽管主动学习可提升样本效率,现有方法多采用G-或D-最优等经典实验设计准则,这些目标未针对偏好学习结构优化,导致算法设计不匹配。本文识别出偏好学习中一个关键直觉,质疑现有准则的适用性。据此提出两种主动学习算法:其一首次提供该场景下的实例相关标签复杂度保证;其二为简单实用的贪心策略。在真实偏好数据集上的评估显示,所提方法相比现有方法显著提升样本效率。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) depends on high-quality datasets of human preference labels, which are costly to collect. Although active learning has been studied to improve sample efficiency relative to passive collection, many existing approaches adopt classical experimental design criteria such as G- or D-optimality. These objectives are not tailored to the structure of preference learning, leaving open the design of problem-specific algorithms. In this work, we identify a simple intuition specific to preference learning that calls into question the suitability of these existing design objectives. Motivated by this insight, we propose two active learning algorithms. The first provides the first instance-dependent label complexity guarantee for this setting, and the second is a simple, practical greedy method. We evaluate our algorithm on real-world preference datasets and observe improved sample efficiency compared to existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。