arXiv:2505.12754cs.LG2025-05被引 1

基于人类偏好筛选指令数据,提升模型生成多样性与满意度。

ProDS: Preference-oriented Data Selection for Instruction Tuning

  • 用直接偏好优化捕捉人类对不同回答的偏好
  • 结合正负偏好评分,选出更符合用户偏好的训练样本
  • 适合需要高质量指令微调数据的场景

指令数据筛选旨在从训练集中选出一个高质量子集,在目标任务上表现不逊于全量数据。现有方法关注指令-响应映射关系,但忽视了人类对多样化回答的偏好。本文提出面向偏好的数据选择方法(ProDS),根据训练样本与目标集观察到的人类偏好的一致性进行评分。核心创新在于将数据选择标准从单纯估计准确响应特征,转向显式对齐目标任务中的人类偏好。具体而言,采用直接偏好优化(DPO)估计多样回答中的偏好;设计双向偏好合成策略,综合正向与负向偏好评分训练样本。大量实验表明,其性能优于现有无任务特异性及目标导向方法。

原文摘要 · Abstract (English)

Instruction data selection aims to identify a high-quality subset from the training set that matches or exceeds the performance of the full dataset on target tasks. Existing methods focus on the instruction-to-response mapping, but neglect the human preference for diverse responses. In this paper, we propose Preference-oriented Data Selection method (ProDS) that scores training samples based on their alignment with preferences observed in the target set. Our key innovation lies in shifting the data selection criteria from merely estimating features for accurate response generation to explicitly aligning training samples with human preferences in target tasks. Specifically, direct preference optimization (DPO) is employed to estimate human preferences across diverse responses. Besides, a bidirectional preference synthesis strategy is designed to score training samples according to both positive preferences and negative preferences. Extensive experimental results demonstrate our superiority to existing task-agnostic and targeted methods.

数据筛选指令微调偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。