提出激励机制,让私有数据持有者安全协作训练个性化模型。
Incentivizing Inclusive Contributions in Model Sharing Markets
- 构建基于图优化的模型共享市场,用博弈论设计激励机制。
- 在11项AI任务中经济收益最高,模型性能优于或相当基线方法。
- 适合关注隐私保护与数据协作的开发者与研究者。
尽管数据在训练现代AI模型中至关重要,但公开可用的数据预计将在几年内耗尽,促使全球关注大规模分散的私有数据。然而,原始数据的敏感性及缺乏激励机制阻碍了这些数据的充分利用。为此,本文提出包容性且可激励的个性化联邦学习(iPFL),通过不暴露原始数据的方式,激励具有不同目标的数据持有者协同训练个性化模型。iPFL通过求解基于图的训练优化问题构建模型共享市场,并采用基于博弈论的激励机制。理论分析表明,iPFL满足个体理性与真实性两个关键激励性质。在11项AI任务(如大语言模型的指令遵循任务)上的实证研究显示,iPFL在经济效用上始终领先,且模型性能优于或相当基线方法。我们预期iPFL将成为未来利用分散私有数据提升AI模型的重要技术,使各方皆满意。
原文摘要 · Abstract (English)
While data plays a crucial role in training contemporary AI models, it is acknowledged that valuable public data will be exhausted in a few years, directing the world's attention towards the massive decentralized private data. However, the privacy-sensitive nature of raw data and lack of incentive mechanism prevent these valuable data from being fully exploited. Addressing these challenges, this paper proposes inclusive and incentivized personalized federated learning (iPFL), which incentivizes data holders with diverse purposes to collaboratively train personalized models without revealing raw data. iPFL constructs a model-sharing market by solving a graph-based training optimization and incorporates an incentive mechanism based on game theory principles. Theoretical analysis shows that iPFL adheres to two key incentive properties: individual rationality and truthfulness. Empirical studies on eleven AI tasks (e.g., large language models' instruction-following tasks) demonstrate that iPFL consistently achieves the highest economic utility, and better or comparable model performance compared to baseline methods. We anticipate that our iPFL can serve as a valuable technique for boosting future AI models on decentralized private data while making everyone satisfied.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。