用帕累托高质量数据提升多目标对齐效率
ParetoHqD: Fast Offline Multiobjective Alignment of Large Language Models using Pareto High-quality Data
- 将人类偏好转化为目标空间中的方向,筛选帕累托前沿附近高质量数据
- 两阶段微调,每阶段使用匹配偏好方向的专属高质量数据集
- 在两个任务上优于五种基线,适合需要多维度对齐的模型优化
将大型语言模型与多种人类期望和价值观对齐,对于满足多样化用户需求至关重要。现有的离线多目标对齐算法(如Rewards-in-Context)虽表现良好且高效,但不当的偏好表示方式以及奖励分数不平衡的训练会限制其性能。本文提出ParetoHqD,通过将人类偏好表示为目标空间中的偏好方向,并将接近帕累托前沿的数据视为“高质量”数据来解决上述问题。针对每个偏好,ParetoHqD采用两阶段监督微调流程,每阶段使用最匹配该偏好方向的专属帕累托高质量数据集。实验结果表明,ParetoHqD在两个多目标对齐任务上均显著优于五种基线方法。
原文摘要 · Abstract (English)
Aligning large language models with multiple human expectations and values is crucial for ensuring that they adequately serve a variety of user needs. To this end, offline multiobjective alignment algorithms such as the Rewards-in-Context algorithm have shown strong performance and efficiency. However, inappropriate preference representations and training with imbalanced reward scores limit the performance of such algorithms. In this work, we introduce ParetoHqD that addresses the above issues by representing human preferences as preference directions in the objective space and regarding data near the Pareto front as "high-quality" data. For each preference, ParetoHqD follows a two-stage supervised fine-tuning process, where each stage uses an individual Pareto high-quality training set that best matches its preference direction. The experimental results have demonstrated the superiority of ParetoHqD over five baselines on two multiobjective alignment tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。