arXiv:2410.10093cs.CLcs.LG2024-10EMNLP被引 18

用自模仿学习高效对齐大模型与演示数据,无需复杂对抗训练。

How to Leverage Demonstration Data in Alignment for Large Language Model? A Self-Imitation Learning Perspective

  • 基于密度比估计构建替代目标,用分类损失优化模仿学习。
  • 在HumanEval、GSM8K等任务上显著优于基线,提升代码生成与数学推理能力。
  • 适合需要轻量微调的模型对齐场景,尤其适用于离线演示数据。

本文提出一种新型广义自模仿学习(GSIL)框架,能高效对齐大语言模型与离线演示数据。通过密度比估计推导出模仿学习的代理目标,使自生成数据得以有效利用,并以简单分类损失优化模仿目标。该方法避免了标准模仿学习中复杂的对抗训练,实现轻量化、高效的模型微调。此外,GSIL包含一组由凸函数参数化的离线损失,提供统一视角以实现演示数据对齐。大量实验表明,GSIL在多个挑战性基准测试中持续显著优于基线,包括代码生成(HumanEval)、数学推理(GSM8K)和指令遵循(MT-Bench)。

原文摘要 · Abstract (English)

This paper introduces a novel generalized self-imitation learning ($\textbf{GSIL}$) framework, which effectively and efficiently aligns large language models with offline demonstration data. We develop $\textbf{GSIL}$ by deriving a surrogate objective of imitation learning with density ratio estimates, facilitating the use of self-generated data and optimizing the imitation learning objective with simple classification losses. $\textbf{GSIL}$ eliminates the need for complex adversarial training in standard imitation learning, achieving lightweight and efficient fine-tuning for large language models. In addition, $\textbf{GSIL}$ encompasses a family of offline losses parameterized by a general class of convex functions for density ratio estimation and enables a unified view for alignment with demonstration data. Extensive experiments show that $\textbf{GSIL}$ consistently and significantly outperforms baselines in many challenging benchmarks, such as coding (HuamnEval), mathematical reasoning (GSM8K) and instruction-following benchmark (MT-Bench).

自模仿学习大模型对齐离线数据轻量微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。