构建千亿级数据共享平台,实现模型训练数据的透明分成。
1 Trillion Token (1TT) Platform: A Novel Framework for Efficient Data Sharing and Compensation in Large Language Models

- 通过透明收益分成机制激励数据贡献者共享非公开数据。
- 数据消费者按约定比例分享服务收入给贡献者,形成可持续协作。
- 适合关注数据合规、模型训练生态建设的研究者与企业。
本文提出1万亿令牌平台(1TT Platform),一种新型框架,旨在促进高效的数据共享并建立透明公平的收益分配机制。该平台促成数据贡献者(提供非公开数据集)与数据使用者(用于提升自身服务)之间的协作。数据贡献者将获得货币补偿,其收入来自数据使用者服务所产生的收益分成。数据使用者需根据预设的收益分配协议,向贡献者共享部分收入。通过引入透明的收益分配模式,1TT平台激励大规模数据共享,推动自然语言处理与大语言模型技术的发展。
原文摘要 · Abstract (English)
In this paper, we propose the 1 Trillion Token Platform (1TT Platform), a novel framework designed to facilitate efficient data sharing with a transparent and equitable profit-sharing mechanism. The platform fosters collaboration between data contributors, who provide otherwise non-disclosed datasets, and a data consumer, who utilizes these datasets to enhance their own services. Data contributors are compensated in monetary terms, receiving a share of the revenue generated by the services of the data consumer. The data consumer is committed to sharing a portion of the revenue with contributors, according to predefined profit-sharing arrangements. By incorporating a transparent profit-sharing paradigm to incentivize large-scale data sharing, the 1TT Platform creates a collaborative environment to drive the advancement of NLP and LLM technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。