用平滑切比雪夫标量化解决多目标强化学习的非凸优化难题。
Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization

- 将多目标强化学习转化为可标量化的优化问题,采用平滑切比雪夫方法避免线性标量化缺陷。
- 在九个蛋白工程任务中,八项表现超越当前最优基线,提升超体积指标。
- 适合需要同时优化多个冲突目标的场景,如蛋白质设计与对话系统对齐。
大型语言模型可通过小规模标注数据集上的离线强化学习实现与人类偏好的对齐。尽管单目标对齐已有深入研究,但现实应用常需同时优化多个冲突奖励,例如蛋白质工程中的催化活性与特异性,或聊天机器人中的帮助性与无害性。以往工作多依赖线性奖励标量化,但该方法无法恢复帕累托前沿的非凸区域。本文提出一种新思路:不直接标量化奖励,而是将多目标强化学习本身视为一个需标量化的优化问题,采用近期提出的平滑切比雪夫标量化技术,克服线性标量化局限。基于此,我们提出光滑切比雪夫多目标偏好优化算法(STOMP),通过标准化各奖励的观测分布,将直接偏好优化推广至多目标设置。我们在三个自回归蛋白语言模型上,利用三个实验室蛋白适应度数据集进行实验验证。相比现有最先进基线,STOMP在八种设置下均取得最高超体积,涵盖离线离策略评估与生成式评估。结果表明,STOMP是一种强大且鲁棒的多目标对齐算法,可显著提升后训练模型在多属性蛋白优化及其他任务中的性能。
原文摘要 · Abstract (English)
Large language models can be aligned with human preferences through offline reinforcement learning (RL) on small labeled datasets. While single-objective alignment is well-studied, many real-world applications demand the simultaneous optimization of multiple conflicting rewards, e.g. optimizing both catalytic activity and specificity in protein engineering, or helpfulness and harmlessness for chatbots. Prior work has largely relied on linear reward scalarization, but this approach provably fails to recover non-convex regions of the Pareto front. In this paper, instead of scalarizing the rewards directly, we frame multi-objective RL itself as an optimization problem to be scalarized via smooth Tchebysheff scalarization, a recent technique that overcomes the shortcomings of linear scalarization. We use this formulation to derive Smooth Tchebysheff Optimization of Multi-Objective Preferences (STOMP), a novel offline RL algorithm that extends direct preference optimization to the multi-objective setting in a principled way by standardizing the individual rewards based on their observed distributions. We empirically validate STOMP on a range of protein engineering tasks by aligning three autoregressive protein language models on three laboratory datasets of protein fitness. Compared to state-of-the-art baselines, STOMP achieves the highest hypervolumes in eight of nine settings according to both offline off-policy and generative evaluations. We thus demonstrate that STOMP is a powerful, robust multi-objective alignment algorithm that can meaningfully improve post-trained models for multi-attribute protein optimization and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。