用协作博弈生成高质量偏好数据,提升大模型对齐效果。
Anyprefer: An Agentic Framework for Preference Data Synthesis
- 构建目标模型与判别模型的协作博弈框架,引入外部工具辅助评分。
- 生成58000条偏好数据,多任务平均提升18.55%至30.05%。
- 适合需要高质量偏好数据的模型对齐研究者使用。
高质量偏好数据对通过偏好学习将基础模型与人类价值观对齐至关重要。然而,人工标注成本高且耗时。现有自奖励方法中,目标模型生成并标注自身响应,因奖励模型与目标模型共享参数,易放大固有偏差。为此,我们提出Anyprefer框架,将数据合成视为一个合作型双玩家马尔可夫博弈,目标模型与判别模型协同工作。引入一系列外部工具辅助判别模型准确评估目标模型输出,缓解评分偏差。同时,设计反馈机制优化双方提示词,增强协作,提升数据质量。合成数据形成新偏好数据集Anyprefer-V1,包含58,000条高质量偏好对。大量实验表明,Anyprefer在四大应用方向、21个数据集上显著提升模型对齐性能:自然语言生成任务平均提升18.55%(5个数据集),视觉-语言理解任务提升3.66%(9个数据集),医学图像分析任务提升30.05%(3个数据集),视觉-运动控制任务提升16.00%(4个任务)。
原文摘要 · Abstract (English)
High-quality preference data is essential for aligning foundation models with human values through preference learning. However, manual annotation of such data is often time-consuming and costly. Recent methods often adopt a self-rewarding approach, where the target model generates and annotates its own preference data, but this can lead to inaccuracies since the reward model shares weights with the target model, thereby amplifying inherent biases. To address these issues, we propose Anyprefer, a framework designed to synthesize high-quality preference data for aligning the target model. Anyprefer frames the data synthesis process as a cooperative two-player Markov Game, where the target model and the judge model collaborate together. Here, a series of external tools are introduced to assist the judge model in accurately rewarding the target model's responses, mitigating biases in the rewarding process. In addition, a feedback mechanism is introduced to optimize prompts for both models, enhancing collaboration and improving data quality. The synthesized data is compiled into a new preference dataset, Anyprefer-V1, consisting of 58K high-quality preference pairs. Extensive experiments show that Anyprefer significantly improves model alignment performance across four main applications, covering 21 datasets, achieving average improvements of 18.55% in five natural language generation datasets, 3.66% in nine vision-language understanding datasets, 30.05% in three medical image analysis datasets, and 16.00% in four visuo-motor control tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。