arXiv:2505.08507cs.LG2025-05NAACL被引 8

InfoPO通过最大化互信息实现大模型对齐,提升推理任务表现

InfoPO: On Mutual Information Maximization for Large Language Model Alignment

  • 基于互信息最大化,无需依赖布拉德利-特里模型
  • 在多个基准上优于现有方法,推理任务提升显著
  • 适合追求高效对齐与强推理能力的研究者

我们研究使用人类偏好数据对大型语言模型进行后训练。近期的直接偏好优化及其变体在模型对齐方面展现出巨大潜力,无需奖励模型和在线采样。然而,这些方法依赖于布拉德利-特里(BT)模型的显式假设,容易过拟合,尤其在推理密集型任务中表现不佳。为此,我们提出一种原理性偏好微调算法InfoPO,能有效且高效地利用偏好数据对齐大模型。InfoPO摒弃了对BT模型的依赖,并防止被选回应的似然度下降。大量实验表明,InfoPO在广泛使用的开源基准上持续优于现有基线,尤其在推理任务中表现突出。

原文摘要 · Abstract (English)

We study the post-training of large language models (LLMs) with human preference data. Recently, direct preference optimization and its variants have shown considerable promise in aligning language models, eliminating the need for reward models and online sampling. Despite these benefits, these methods rely on explicit assumptions about the Bradley-Terry (BT) model, which makes them prone to overfitting and results in suboptimal performance, particularly on reasoning-heavy tasks. To address these challenges, we propose a principled preference fine-tuning algorithm called InfoPO, which effectively and efficiently aligns large language models using preference data. InfoPO eliminates the reliance on the BT model and prevents the likelihood of the chosen response from decreasing. Extensive experiments confirm that InfoPO consistently outperforms established baselines on widely used open benchmarks, particularly in reasoning tasks.

大模型对齐偏好优化互信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。