arXiv:2501.03271cs.LGcs.AI2025-01ACL被引 2

用核方法增强偏好优化,让大模型更懂人类多样需求

DPO Kernels: A Semantically-Aware, Kernel-Enhanced, and Divergence-Rich Paradigm for Direct Preference Optimization

  • 引入多种核函数实现更丰富的特征变换
  • 在12个数据集上实现事实性与安全性最优表现
  • 自动选择最佳核-散度组合,适合对齐研究者使用

大语言模型的快速发展带来了诸多应用,但也凸显了与多样化价值观对齐的挑战。直接偏好优化(DPO)虽为核心方法,却受限于固定散度和有限特征变换。本文提出DPO-Kernels,通过四个关键贡献解决此问题:(i) 使用多项式、RBF、马哈拉诺比斯和谱核构建核化表示,并结合基于嵌入与概率的混合损失;(ii) 引入杰恩森-香农、海林格、瑞尼、巴塔查里亚、沃尔瑟斯坦及f-散度等多种散度以提升稳定性;(iii) 设计数据驱动的选择指标,自动优选最佳核-散度组合;(iv) 采用分层核混合机制兼顾局部精度与全局建模。在12个数据集上的评估显示,该方法在事实性、安全性、推理能力与指令遵循方面均达到领先水平。基于重尾自正则化理论,DPO-Kernels保持了强泛化能力,为后续对齐研究提供全面工具。

原文摘要 · Abstract (English)

The rapid rise of large language models (LLMs) has unlocked many applications but also underscores the challenge of aligning them with diverse values and preferences. Direct Preference Optimization (DPO) is central to alignment but constrained by fixed divergences and limited feature transformations. We propose DPO-Kernels, which integrates kernel methods to address these issues through four key contributions: (i) Kernelized Representations with polynomial, RBF, Mahalanobis, and spectral kernels for richer transformations, plus a hybrid loss combining embedding-based and probability-based objectives; (ii) Divergence Alternatives (Jensen-Shannon, Hellinger, Renyi, Bhattacharyya, Wasserstein, and f-divergences) for greater stability; (iii) Data-Driven Selection metrics that automatically choose the best kernel-divergence pair; and (iv) a Hierarchical Mixture of Kernels for both local precision and global modeling. Evaluations on 12 datasets demonstrate state-of-the-art performance in factuality, safety, reasoning, and instruction following. Grounded in Heavy-Tailed Self-Regularization, DPO-Kernels maintains robust generalization for LLMs, offering a comprehensive resource for further alignment research.

大模型对齐偏好优化核方法强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。