arXiv:2506.07272cs.LG2025-06NeurIPS被引 2

用新统计方法激励数据共享者说真话,防造假、提质量。

A Cramér-von Mises Approach to Incentivizing Truthful Data Sharing

  • 基于Cramér-von Mises双样本检验设计激励机制,不依赖高斯假设。
  • 理论证明说真话是纳什均衡,真实数据越多奖励越高。
  • 在语言和图像数据上验证有效,适合数据市场与联盟场景。

现代数据市场和数据共享联盟越来越依赖激励机制来鼓励参与者贡献数据。然而,基于提交数据量的奖励机制易被操纵,参与者可能提交伪造或低质量数据以虚增收益。以往工作通过比较各参与者数据来促进诚实:当他人提供真实数据时,最小化差异的最佳方式就是自己也提交真实数据。但先前方法依赖强假设(如数据服从高斯分布),限制了应用范围。本文提出一种基于新型双样本检验的奖励机制,灵感来自Cramér-von Mises统计量。该方法严格激励参与者提交更多真实数据,同时抑制数据伪造及其他不诚实行为。我们证明,在贝叶斯与先验无关两种设定下,诚实报告构成(近似)纳什均衡。理论上将方法应用于三个典型数据共享问题,显著放宽了先前工作的关键假设。实证方面,通过模拟及真实语言与图像数据验证了机制的有效性。

原文摘要 · Abstract (English)

Modern data marketplaces and data sharing consortia increasingly rely on incentive mechanisms to encourage agents to contribute data. However, schemes that reward agents based on the quantity of submitted data are vulnerable to manipulation, as agents may submit fabricated or low-quality data to inflate their rewards. Prior work has proposed comparing each agent's data against others' to promote honesty: when others contribute genuine data, the best way to minimize discrepancy is to do the same. Yet prior implementations of this idea rely on very strong assumptions about the data distribution (e.g. Gaussian), limiting their applicability. In this work, we develop reward mechanisms based on a novel, two-sample test inspired by the Cramér-von Mises statistic. Our methods strictly incentivize agents to submit more genuine data, while disincentivizing data fabrication and other types of untruthful reporting. We establish that truthful reporting constitutes a (possibly approximate) Nash equilibrium in both Bayesian and prior-agnostic settings. We theoretically instantiate our method in three canonical data sharing problems and show that it relaxes key assumptions made by prior work. Empirically, we demonstrate that our mechanism incentivizes truthful data sharing via simulations and on real-world language and image data.

数据共享激励机制统计检验真实性验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。