提出首个确保数据真实与公平激励的贝叶斯学习机制
Incentivizing Truthfulness and Collaborative Fairness in Bayesian Learning

- 结合半值与未知验证集的估值函数,保障公平性
- 理论证明在均衡下提交真实数据可最大化收益
- 适用于数据可信度关键的联邦学习场景
协同机器学习通过整合多方数据训练高质量模型。现有数据估值方法虽能公平分配奖励,但无法验证数据真实性,导致数据提供方可通过重复或噪声数据操纵估值以获取更高回报。本文首次提出一个可证明在均衡状态下同时保证(F)协同公平性与(T)数据真实性的机制。该机制融合半值(如Shapley值)与基于未知验证集的数据估值函数(DVF)。由于半值受其他数据影响,我们引入额外条件,证明参与者通过提交真实知识数据可在联盟和半值中实现期望估值最大化。此外,我们探讨了中介预算有限或缺乏验证集时(F)与(T)的合理放松。理论结果在合成与真实数据集上得到验证。
原文摘要 · Abstract (English)
Collaborative machine learning involves training high-quality models using datasets from a number of sources. To incentivize sources to share data, existing data valuation methods fairly reward each source based on its data submitted as is. However, as these methods do not verify nor incentivize data truthfulness, the sources can manipulate their data (e.g., by submitting duplicated or noisy data) to artificially increase their valuations and rewards or prevent others from benefiting. This paper presents the first mechanism that provably ensures (F) collaborative fairness and incentivizes (T) truthfulness at equilibrium for Bayesian models. Our mechanism combines semivalues (e.g., Shapley value), which ensure fairness, and a truthful data valuation function (DVF) based on a validation set that is unknown to the sources. As semivalues are influenced by others' data, we introduce an additional condition to prove that a source can maximize its expected data values in coalitions and semivalues by submitting a dataset that captures its true knowledge. Additionally, we discuss the implications and suitable relaxations of (F) and (T) when the mediator has a limited budget for rewards or lacks a validation set. Our theoretical findings are validated on synthetic and real-world datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。