arXiv:2506.22578cs.LGcs.AI2025-06被引 8

揭示强化学习与对比学习的深层联系,提出新优化方法提升模型推理能力

The Hidden Link Between RLHF and Contrastive Learning

  • 从互信息最大化视角统一解释RLHF与DPO,二者本质是基于基模型样本的对比学习
  • 提出MIO方法,解决DPO后期选中概率下降问题,在数学推理等任务上表现更优
  • 适合对大模型对齐机制、优化算法改进感兴趣的研究者与工程师

大型语言模型(LLMs)与人类价值观对齐近年备受关注,代表性方法包括成本高昂的基于人类反馈的强化学习(RLHF)和简单的直接偏好优化(DPO)。本文证明,两者均可从互信息(MI)最大化角度解读,揭示其与对比学习的深刻关联。在此框架下,RLHF与DPO均被视为基于基模型生成正负样本的对比学习方法,利用Donsker-Varadhan(DV)下界(等价于MINE估计器)。该范式进一步说明,为何RLHF无法内在激励超出基模型已有能力的推理能力。基于此,我们用Jensen-Shannon(JS)MI估计器替代DV/MINE边界,提出互信息优化(MIO)。理论分析与大量实证评估表明,MIO缓解了DPO中观测到的后期选中概率下降现象,在多个具有挑战性的推理与数学基准测试中达到或超越现有性能。

原文摘要 · Abstract (English)

Alignment of large language models (LLMs) with human values has recently garnered significant attention, with prominent examples including the canonical yet costly Reinforcement Learning from Human Feedback (RLHF) and the simple Direct Preference Optimization (DPO). In this work, we demonstrate that both RLHF and DPO can be interpreted from the perspective of mutual information (MI) maximization, uncovering a profound connection to contrastive learning. Within this framework, both RLHF and DPO can be interpreted as methods that performing contrastive learning based on the positive and negative samples derived from base model, leveraging the Donsker-Varadhan (DV) lower bound on MI (equivalently, the MINE estimator). Such paradigm further illuminates why RLHF may not intrinsically incentivize reasoning capacities in LLMs beyond what is already present in the base model. Building on the perspective, we replace the DV/MINE bound with the Jensen-Shannon (JS) MI estimator and propose the Mutual Information Optimization (MIO). Comprehensive theoretical analysis and extensive empirical evaluations demonstrate that MIO mitigates the late-stage decline in chosen-likelihood observed in DPO, achieving competitive or superior performance across various challenging reasoning and mathematical benchmarks.

大模型对齐对比学习优化方法互信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。