arXiv:2411.10914cs.CL2024-11NAACL被引 22

平衡知识广度与深度,提升大模型对齐效果

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment

  • 提出动态增强每样本知识深度的BPO方法
  • 在多个基准上优于基线,且训练效率不降
  • 适合关注对齐数据优化的研究者

近年来,基于人类反馈的强化学习(RLHF)是大型语言模型成功的关键。本文首次引入知识广度与深度的概念,分别衡量模型或知识源的全面性与深入程度。研究发现,提示与回复数量的不平衡会导致对齐数据集中广度与深度学习的失衡;即使采用简单的均匀化方法平衡指令与回复数量,也能带来显著提升。基于此,本文提出平衡偏好优化(BPO),通过梯度聚类动态增强每个样本的知识深度。BPO依据模型优化方向评估每个增强样本的知识信息量与有用性,实现按需学习深度知识。实验结果表明,BPO在多个基准测试中优于其他基线方法,同时保持训练效率。我们还对BPO各组件进行详细分析,为未来偏好数据优化研究提供指导。

原文摘要 · Abstract (English)

Reinforcement Learning with Human Feedback (RLHF) is the key to the success of large language models (LLMs) in recent years. In this work, we first introduce the concepts of knowledge breadth and knowledge depth, which measure the comprehensiveness and depth of an LLM or knowledge source respectively. We reveal that the imbalance in the number of prompts and responses can lead to a potential disparity in breadth and depth learning within alignment tuning datasets by showing that even a simple uniform method for balancing the number of instructions and responses can lead to significant improvements. Building on this, we further propose Balanced Preference Optimization (BPO), designed to dynamically augment the knowledge depth of each sample. BPO is motivated by the observation that the usefulness of knowledge varies across samples, necessitating tailored learning of knowledge depth. To achieve this, we introduce gradient-based clustering, estimating the knowledge informativeness and usefulness of each augmented sample based on the model's optimization direction. Our experimental results across various benchmarks demonstrate that BPO outperforms other baseline methods in alignment tuning while maintaining training efficiency. Furthermore, we conduct a detailed analysis of each component of BPO, providing guidelines for future research in preference data optimization.

大模型对齐偏好优化知识深度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。