arXiv:2412.06597cs.LG2024-12

让自私的参与者在协作中主动分享数据,通过动态模型激励提升贡献质量。

Self-Interested Agents in Collaborative Machine Learning: An Incentivized Adaptive Data-Centric Framework

  • 用可学习的随机策略让每个参与方自主选择分享数据子集。
  • 引入带权重的损失函数与噪声机制,使各参与方获得差异化模型。
  • 理论保证优化过程收敛,适合存在数据偏移的分布式学习场景。

我们提出一种面向自利参与方的自适应数据中心协作机器学习框架,由仲裁者协调。该框架支持在线增量学习:每轮中,仲裁者从参与方收集数据批次,训练模型,并为每个参与方生成反映其数据贡献的独特模型。这一设计形成反馈循环:共享数据影响模型更新,而模型表现又指导未来的数据共享策略。参与方基于自身评估函数,通过参数化随机策略评估并划分数据,利用策略梯度方法优化以最大化所获模型效用。仲裁者侧则优化真实数据分布上的期望损失函数,引入参与方特定权重以应对异源数据和选择性共享带来的分布差异。采用双层优化算法联合学习模型参数与参与方权重。通过畸变函数计算的均值为零噪声,用于调整权重并生成差异化模型,促进有价值的数据共享,无需单独训练。框架具备非渐近分析基础,确保参与方策略优化收敛至评估函数的近似驻点,仲裁者优化收敛至期望损失函数的近似驻点。

原文摘要 · Abstract (English)

We propose a framework for adaptive data-centric collaborative machine learning among self-interested agents, coordinated by an arbiter. Designed to handle the incremental nature of real-world data, the framework operates in an online manner: at each time step, the arbiter collects a batch of data from agents, trains a machine learning model, and provides each agent with a distinct model reflecting its data contributions. This setup establishes a feedback loop where shared data influence model updates, and the resulting models guide future data-sharing policies. Agents evaluate and partition their data, selecting a partition to share using a stochastic parameterized policy, learned via policy gradient methods to optimize the utility of the received model as defined by agent-specific evaluation functions. On the arbiter side, the expected loss function over the true data distribution is optimized, incorporating agent-specific weights to account for distributional differences arising from diverse sources and selective sharing. A bilevel optimization algorithm jointly learns the model parameters and agent-specific weights. Mean-zero noise, computed using a distortion function that adjusts these agent-specific weights, is introduced to generate distinct agent-specific models, promoting valuable data sharing without requiring separate training. Our framework is underpinned by non-asymptotic analyses, ensuring convergence of the agent-side policy optimization to an approximate stationary point of the evaluation functions and convergence of the arbiter-side optimization to an approximate stationary point of the expected loss function.

协同学习激励机制数据共享双层优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。