为大模型训练设计公平可信的数据共享机制,兼顾质量与激励。
Designing DSIC Mechanisms for Data Sharing in the Era of Large Language Models
- 基于数据质量和贡献度设计虚拟成本排名与支付机制
- 在预算有限下仍能获取更高质量数据,且防谎防合谋
- 适合构建可持续的大模型数据交易市场,尤其关注隐私与公平
训练大语言模型需要大量高质量数据,但机构面临法律、隐私和战略限制。现有数据采购方式常依赖不可验证的信任或忽略提供者成本差异。本文提出一个机制设计框架,实现主导策略激励相容(DSIC)、个体理性与弱预算平衡,根据数据质量与学习效用进行奖励。模型假设提供者私有知晓自身数据成本与质量,价值仅来自对模型性能的贡献。提出质量加权边际激励拍卖(Q-MIA),通过虚拟成本排序与Myerson式支付确保DSIC与预算可行性。针对流动性不足或长期激励场景,引入边际效用代币(MUT),按边际贡献分配未来权益。混合机制Mixed-MIA融合即时支付与延期奖励。所有机制支持可验证、隐私保护实现。理论与实证均表明,其优于基于数量或信任的基线方法,在预算约束下能有效获取高质量数据,并对误报与共谋保持鲁棒性。为未来大模型可持续、公平的数据市场奠定理论基础。
原文摘要 · Abstract (English)
Training large language models (LLMs) requires vast amounts of high-quality data from institutions that face legal, privacy, and strategic constraints. Existing data procurement methods often rely on unverifiable trust or ignore heterogeneous provider costs. We introduce a mechanism-design framework for truthful, trust-minimized data sharing that ensures dominant-strategy incentive compatibility (DSIC), individual rationality, and weak budget balance, while rewarding data based on both quality and learning utility. We formalize a model where providers privately know their data cost and quality, and value arises solely from the data's contribution to model performance. Based on this, we propose the Quality-Weighted Marginal-Incentive Auction (Q-MIA), which ranks providers using a virtual cost metric and uses Myerson-style payments to ensure DSIC and budget feasibility. To support settings with limited liquidity or long-term incentives, we introduce the Marginal Utility Token (MUT), which allocates future rights based on marginal contributions. We unify these in Mixed-MIA, a hybrid mechanism balancing upfront payments and deferred rewards. All mechanisms support verifiable, privacy-preserving implementation. Theoretically and empirically, they outperform volume-based and trust-based baselines, eliciting higher-quality data under budget constraints while remaining robust to misreporting and collusion. This establishes a principled foundation for sustainable and fair data markets for future LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。