arXiv:2506.04063cs.HCcs.DC2025-06被引 2

用众包方式提升大模型对齐效果,更公平高效。

Crowd-SFT: Crowdsourcing for LLM Alignment

  • 引入积分奖励机制,按贡献度分配激励
  • 多模型选择使目标距离减少55%
  • 适合需要大规模、公平标注的场景

大型语言模型(LLMs)越来越多地依赖监督微调(SFT)和基于人类反馈的强化学习(RLHF)来对齐模型输出与人类偏好。虽然RLHF采用独立的奖励模型进行强化学习,而SFT则使用人工标注的数据集进行监督学习,但两者传统上依赖小规模、经过筛选的标注员,导致成本高、易产生偏见且难以扩展。我们提出一种开放的众包微调框架,通过广泛收集反馈来改进SFT,无需大量标注培训。该框架通过与谢林值(Shapley values)相关的积分奖励系统实现激励公平,并通过迭代模型更新引导收敛。多模型选择框架相比单模型选择,目标距离降低高达55%。后续实验验证了积分奖励机制与谢林值高度一致,支持了公平且可扩展的参与。

原文摘要 · Abstract (English)

Large Language Models (LLMs) increasingly rely on Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF) to align model responses with human preferences. While RLHF employs a reinforcement learning approach with a separate reward model, SFT uses human-curated datasets for supervised learning. Both approaches traditionally depend on small, vetted groups of annotators, making them costly, prone to bias, and limited in scalability. We propose an open, crowd-sourced fine-tuning framework that addresses these limitations by enabling broader feedback collection for SFT without extensive annotator training. Our framework promotes incentive fairness via a point-based reward system correlated with Shapley values and guides model convergence through iterative model updates. Our multi-model selection framework demonstrates up to a 55% reduction in target distance over single-model selection, enabling subsequent experiments that validate our point-based reward mechanism's close alignment with Shapley values (a well-established method for attributing individual contributions) thereby supporting fair and scalable participation.

大模型对齐众包奖励机制微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。