解决联邦学习中合作与竞争的矛盾,用生成数据策略平衡收益与风险。
Cooperate to Compete: Strategic Data Generation and Incentivization Framework for Coopetitive Cross-Silo Federated Learning

- 将每轮训练建模为权衡性能、成本与竞争损失的博弈决策。
- 在非独立同分布数据下提升社会福利,优于现有基线方法。
- 适合医疗等敏感领域跨机构协作,兼顾隐私与长期合作。
在医疗等数据敏感领域,跨孤岛联邦学习(CFL)允许机构在不共享原始数据的情况下协同训练模型。然而,实际部署中存在合作与竞争并存的矛盾:机构在训练中合作,却在下游市场竞争。此时,数据量、质量与多样性虽能提升全局模型,但也会无意中增强竞争对手。这一困境在非独立同分布(non-IID)数据下加剧,导致学习收益不对称,削弱持续参与意愿。现有竞争感知联邦学习与激励设计方法仅基于边际贡献奖励,未考虑强化对手的成本。本文提出CoCoGen+,一个兼容合谋竞争的数据生成与激励框架,联合建模非IID数据与组织间竞争,并将生成式AI合成数据作为战略性决策。具体而言,将每轮训练设为加权势博弈,机构通过权衡学习收益、计算成本与竞争带来的效用损失,决定合成数据生成量。我们给出可解析的均衡特征,并推导可实施的生成策略以最大化社会福利。为促进长期合作,引入基于收益再分配的激励机制,补偿组织的贡献及竞争带来的效用下降。在多种学习任务上的实验验证了其可行性。结果表明,非IID数据、竞争强度与激励机制共同影响组织策略与社会福利,且CoCoGen+在效率上优于基线。
原文摘要 · Abstract (English)
In data-sensitive domains such as healthcare, cross-silo federated learning (CFL) allows organizations to collaboratively train AI models without sharing raw data. However, practical CFL deployments are inherently coopetitive, in which organizations cooperate during model training while competing in downstream markets. In such settings, training contributions, including data volume, quality, and diversity, can improve the global model yet inadvertently strengthen rivals. This dilemma is amplified by non-IID data, which leads to asymmetric learning gains and undermines sustained participation. While existing competition-aware CFL and incentive-design approaches reward organizations based on marginal training contributions, they fail to account for the costs of strengthening competitors. In this paper, we introduce CoCoGen+, a coopetition-compatible data generation and incentivization framework that jointly models non-IID data and inter-organizational competition while endogenizing GenAI-based synthetic data generation as a strategic decision. Specifically, CoCoGen+ formulates each training round as a weighted potential game, where organizations strategically decide how much synthetic data to generate by balancing learning performance gains against computational costs and competition-caused utility losses. We then provide a tractable equilibrium characterization and derive implementable generation strategies to maximize social welfare. To promote long-term collaboration, we integrate a payoff redistribution-based incentive mechanism to compensate organizations for their contributions and competition-caused utility degradation. Experiments on varying learning tasks validate the feasibility of CoCoGen+. The results show how non-IID data, competition intensity, and incentives shape organizational strategies and social welfare, while CoCoGen+ outperforms baselines in efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。