arXiv:2412.16475cs.LGcs.AI2024-12被引 1

提出代理反馈提升学习效率的条件,指导大模型高效利用有限专家数据

When Can Proxies Improve the Sample Complexity of Preference Learning?

  • 给出代理反馈有效提升样本效率的数学条件
  • 证明在满足条件时可显著减少所需训练数据量
  • 适用于医疗、教育等专家数据稀缺场景

我们研究了奖励欺骗问题——最大化代理奖励未必提升真实奖励。这对大型语言模型(LLMs)至关重要,因它们常基于可能不反映真实目标的人类偏好进行微调。现有方法采用正则化、奖励模型调整及奖励欺骗检测等技巧来限制代理偏好对模型的影响。幸运的是,在医学、教育、法律等领域,通常存在少量专家数据。在此背景下,尚不清楚添加代理数据是否能改善策略学习。本文提出了代理反馈需满足的一组充分条件,若满足,则可保证代理数据能显著降低学习真实策略所需的样本复杂度。这些条件可指导特定任务的数据收集过程。结果表明,可通过参数化方式实现这一改进,并详细说明如何改造现有架构以达成更优样本效率。

原文摘要 · Abstract (English)

We address the problem of reward hacking, where maximising a proxy reward does not necessarily increase the true reward. This is a key concern for Large Language Models (LLMs), as they are often fine-tuned on human preferences that may not accurately reflect a true objective. Existing work uses various tricks such as regularisation, tweaks to the reward model, and reward hacking detectors, to limit the influence that such proxy preferences have on a model. Luckily, in many contexts such as medicine, education, and law, a sparse amount of expert data is often available. In these cases, it is often unclear whether the addition of proxy data can improve policy learning. We outline a set of sufficient conditions on proxy feedback that, if satisfied, indicate that proxy data can provably improve the sample complexity of learning the ground truth policy. These conditions can inform the data collection process for specific tasks. The result implies a parameterisation for LLMs that achieves this improved sample complexity. We detail how one can adapt existing architectures to yield this improved sample complexity.

偏好学习大模型样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。