arXiv:2505.04993cs.CL2025-05ICML被引 7

用离散隐变量建模人类偏好复杂性,提升大模型对齐效果

Latent Preference Coding: Aligning Large Language Models via Discrete Latent Codes

  • 通过离散隐变量捕捉偏好背后的多重因素及其组合
  • 在多个基准上使DPO、SimPO、IPO算法性能全面提升
  • 无需预设奖励函数,适合需要鲁棒对齐的实用场景

大语言模型虽取得显著进展,但其生成结果与人类偏好的对齐仍是关键挑战。现有偏好建模方法多依赖显式或隐式奖励函数,忽视了人类偏好在不同任务和群体中可能存在的复杂且矛盾的特征。为此,本文提出潜变量偏好编码(LPC)框架,利用离散隐变量建模整体偏好背后的隐含因素及其组合。LPC可无缝集成多种离线对齐算法,仅从数据中自动推断潜在因素及其重要性,无需预定义奖励函数或人工设计权重。在多个基准上的实验证明,LPC在三种基础模型(Mistral-7B、Llama3-8B、Llama3-8B-Instruct)上持续提升了DPO、SimPO、IPO三种对齐算法的性能。深入分析显示,学习到的隐变量能有效捕捉人类偏好分布差异,并显著增强对数据噪声的鲁棒性。该方法为构建更稳健、通用的对齐技术提供了统一表征,推动强大大模型的负责任部署。

原文摘要 · Abstract (English)

Large language models (LLMs) have achieved remarkable success, yet aligning their generations with human preferences remains a critical challenge. Existing approaches to preference modeling often rely on an explicit or implicit reward function, overlooking the intricate and multifaceted nature of human preferences that may encompass conflicting factors across diverse tasks and populations. To address this limitation, we introduce Latent Preference Coding (LPC), a novel framework that models the implicit factors as well as their combinations behind holistic preferences using discrete latent codes. LPC seamlessly integrates with various offline alignment algorithms, automatically inferring the underlying factors and their importance from data without relying on pre-defined reward functions and hand-crafted combination weights. Extensive experiments on multiple benchmarks demonstrate that LPC consistently improves upon three alignment algorithms (DPO, SimPO, and IPO) using three base models (Mistral-7B, Llama3-8B, and Llama3-8B-Instruct). Furthermore, deeper analysis reveals that the learned latent codes effectively capture the differences in the distribution of human preferences and significantly enhance the robustness of alignment against noise in data. By providing a unified representation for the multifarious preference factors, LPC paves the way towards developing more robust and versatile alignment techniques for the responsible deployment of powerful LLMs.

大模型对齐偏好建模隐变量编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。