arXiv:2409.18417cs.LGcs.AI2024-09被引 3

用拍卖机制降低人类反馈强化学习的数据成本

VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback

  • 引入拍卖机制优化人类反馈数据收集成本
  • 在保持模型性能前提下显著降低训练开支
  • 适合关注低成本高效微调的工程团队

本文聚焦于人类反馈强化学习(RLHF)中的成本效率问题。尽管人类偏好标注存在直接经济成本,但现有研究未考虑偏好数据集的经济效用。由于偏好数据中常存在复杂非传递性或循环关系,现有微调算法难以全面捕捉偏好,导致生产环境中数据积累带来严重成本效率问题。本文将LLM微调视为一种货币化经济行为,提出基于拍卖机制的偏好数据采集协议。实验表明,该方法在保证模型性能的同时,显著提升了数据使用的经济效率,尤其适用于以高质量反馈为核心的微调任务。

原文摘要 · Abstract (English)

This paper addresses the cost-efficiency aspect of Reinforcement Learning from Human Feedback (RLHF). RLHF leverages datasets of human preferences over outputs of large language models (LLM)s to instill human expectations into LLMs. Although preference annotation comes with a monetized cost, the economic utility of a preference dataset has not been considered by far. What exacerbates this situation is that, given complex intransitive or cyclic relationships in preference datasets, existing algorithms for fine-tuning LLMs are still far from capturing comprehensive preferences. This raises severe cost-efficiency concerns in production environments, where preference data accumulate over time. In this paper, we discuss the fine-tuning of LLMs as a monetized economy and introduce an auction mechanism to improve the efficiency of preference data collection in dollar terms. We show that introducing an auction mechanism can play an essential role in enhancing the cost-efficiency of RLHF, while maintaining satisfactory model performance. Experimental results demonstrate that our proposed auction-based protocol is cost-effective for fine-tuning LLMs concentrating on high-quality feedback.

RLHF数据成本拍卖机制大模型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。