arXiv:2502.00198cs.GTcs.CL2025-02NeurIPS被引 10

为大模型训练数据设计公平定价机制,让标注者获益并提升数据质量。

Fairshare Data Pricing via Data Valuation for Large Language Models

论文配图:Fairshare Data Pricing via Data Valuation for Large Language Models
图 1 · 摘自论文原文
  • 基于数据贡献估值构建公平分享定价机制
  • 实测提升标注者收入与高质量数据供给稳定性
  • 适合关注数据伦理与可持续模型训练的研究者

大语言模型的训练数据是其核心,但当前数据市场常存在剥削性定价——从边缘群体获取数据却极少支付报酬或给予认可。本文提出一个大模型数据市场的理论框架,建模买方(模型开发者)与卖方(人工标注者)的战略互动。理论与实证分析表明,剥削性定价会驱逐高质量卖家,降低数据质量并损害长期模型性能。为此,我们提出 fairshare 定价机制,基于数据估值量化每条数据的贡献,通过激励对齐维持卖家参与度,优化买卖双方效用。理论上,fairshare 实现双赢:最大化买方长期效用与卖方利润,同时维持市场参与。在开源大模型训练复杂NLP任务(如数学问题、医疗诊断、物理推理)的实验中,fairshare 显著提升卖家收益,保障高质量数据持续供给,同时提高买方单位成本性能与长期福利。研究为大模型数据市场提供了公平、透明且经济可持续的实践路径。

原文摘要 · Abstract (English)

Training data is the backbone of large language models (LLMs), yet today's data markets often operate under exploitative pricing -- sourcing data from marginalized groups with little pay or recognition. This paper introduces a theoretical framework for LLM data markets, modeling the strategic interactions between buyers (LLM builders) and sellers (human annotators). We begin with theoretical and empirical analysis showing how exploitative pricing drives high-quality sellers out of the market, degrading data quality and long-term model performance. Then we introduce fairshare, a pricing mechanism grounded in data valuation that quantifies each data's contribution. It aligns incentives by sustaining seller participation and optimizing utility for both buyers and sellers. Theoretically, we show that fairshare yields mutually optimal outcomes: maximizing long-term buyer utility and seller profit while sustaining market participation. Empirically when training open-source LLMs on complex NLP tasks, including math problems, medical diagnosis, and physical reasoning, fairshare boosts seller earnings and ensures a stable supply of high-quality data, while improving buyers' performance-per-dollar and long-term welfare. Our findings offer a concrete path toward fair, transparent, and economically sustainable data markets for LLM.

数据定价公平性大模型训练数据伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。