让大模型同时满足多种人类偏好,通过最大化解集覆盖范围提升多样性与质量。
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
- 采用后验多目标优化思想,通过最大化超体积来生成多样化的模型策略。
- 在无害性、有用性、幽默感等多维度上均优于现有方法,跨数据集表现稳定。
- 适合需要平衡多个复杂人类偏好的场景,如对话系统与内容生成。
大型语言模型的多目标对齐(MOAHF)面临挑战,因人类偏好复杂且常相互冲突。现有研究多基于先验多目标优化,即训练或推理时已知偏好。当偏好未知或难以量化时,更自然的做法是通过多样化解集覆盖帕累托前沿。本文提出算法HaM,旨在学习多样化的语言模型策略以最大化其超体积。这是首个将后验多目标优化应用于MOAHF的工作。HaM计算与空间效率高,在多种数据集上于无害性、有用性、幽默感、忠实性及幻觉控制等多个目标上表现优异。
原文摘要 · Abstract (English)
Multi-objective alignment from human feedback (MOAHF) in large language models (LLMs) is a challenging problem as human preferences are complex, multifaceted, and often conflicting. Recent works on MOAHF considered a-priori multi-objective optimization (MOO), where human preferences are known at training or inference time. In contrast, when human preferences are unknown or difficult to quantify, a natural approach is to cover the Pareto front by multiple diverse solutions. We propose an algorithm HaM for learning diverse LLM policies that maximizes their hypervolume. This is the first application of a-posteriori MOO to MOAHF. HaM is computationally and space efficient, and empirically superior across objectives such as harmlessness, helpfulness, humor, faithfulness, and hallucination, on various datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。