用简单环境发现更优的强化学习对齐算法。
Meta-Learning Objectives for Preference Optimization
- 在简化任务中设计诊断基准,替代昂贵的模型对齐测试。
- 提出镜面偏好优化(MPO)并用进化策略找到适配不同数据的算法。
- 在机器人控制和大模型对齐任务中均显著超越现有方法。
评估大语言模型对齐中的偏好优化(PO)算法面临成本高、噪声大及变量多等挑战。本文证明可在简化基准上获得有效洞察。我们设计了一套基于MuJoCo的任务与数据集,系统性评估PO算法,建立更可控且低成本的基准。随后提出一类基于镜面下降的新型PO算法,称为镜面偏好优化(MPO),并通过进化策略搜索出针对混合质量或含噪数据特化的算法。实验表明,所发现的算法在目标MuJoCo设置中全面优于已有方法。最后,基于这些洞见,设计的新算法在大模型对齐任务中显著超越现有基线。
原文摘要 · Abstract (English)
Evaluating preference optimization (PO) algorithms on LLM alignment is a challenging task that presents prohibitive costs, noise, and several variables like model size and hyper-parameters. In this work, we show that it is possible to gain insights on the efficacy of PO algorithm on simpler benchmarks. We design a diagnostic suite of MuJoCo tasks and datasets, which we use to systematically evaluate PO algorithms, establishing a more controlled and cheaper benchmark. We then propose a novel family of PO algorithms based on mirror descent, which we call Mirror Preference Optimization (MPO). Through evolutionary strategies, we search this class to discover algorithms specialized to specific properties of preference datasets, such as mixed-quality or noisy data. We demonstrate that our discovered PO algorithms outperform all known algorithms in the targeted MuJoCo settings. Finally, based on the insights gained from our MuJoCo experiments, we design a PO algorithm that significantly outperform existing baselines in an LLM alignment task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。