arXiv:2505.21684cs.LGcs.DC2025-05被引 4

构建去中心化LLM训练激励机制,让任何人贡献伪梯度即可获酬。

Incentivizing Permissionless Distributed Learning of LLMs

  • 通过双阶段机制快速筛选节点可用性与可靠性,结合伪梯度贡献前后的损失评估。
  • 在无许可环境下用1.2B参数模型实测,每轮训练均产生可比肩主流的模型性能。
  • 适合关注去中心化训练、激励设计及区块链协同学习的研究者与开发者。

我们提出一种用于基础模型分布式深度学习的激励系统,参与者因贡献而获得奖励。该系统名为Gauntlet,已部署于bittensor区块链,用于训练一个1.2B参数的LLM,完全基于无许可的伪梯度贡献——无需控制注册用户或其硬件。Gauntlet适用于任何依赖更新或伪梯度聚合的同步分布式训练方案。系统采用两阶段机制,快速过滤节点的在线时长、可靠性和同步性,并引入核心组件,估算个体伪梯度贡献前后的损失变化。我们使用OpenSkill评分系统追踪伪梯度得分随时间的竞争力。此外,引入新机制确保网络中节点执行唯一计算。实际运行中,1.2B模型持续支付真实代币给参与者,其每轮迭代表现已达到可竞争水平,验证了该激励系统的有效性。

原文摘要 · Abstract (English)

We describe an incentive system for distributed deep learning of foundational models where peers are rewarded for contributions. The incentive system, \textit{Gauntlet}, has been deployed on the bittensor blockchain and used to train a 1.2B LLM with completely permissionless contributions of pseudo-gradients: no control over the users that can register or their hardware. \textit{Gauntlet} can be applied to any synchronous distributed training scheme that relies on aggregating updates or pseudo-gradients. We rely on a two-stage mechanism for fast filtering of peer uptime, reliability, and synchronization, combined with the core component that estimates the loss before and after individual pseudo-gradient contributions. We utilized an OpenSkill rating system to track competitiveness of pseudo-gradient scores across time. Finally, we introduce a novel mechanism to ensure peers on the network perform unique computations. Our live 1.2B run, which has paid out real-valued tokens to participants based on the value of their contributions, yielded a competitive (on a per-iteration basis) 1.2B model that demonstrates the utility of our incentive system.

去中心化训练激励机制伪梯度LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。