训练数据的人力成本远超模型训练,应获得合理补偿。
Position: The Most Expensive Part of an LLM should be its Training Data
- 估算64个大模型数据集人力成本,按保守工资计算
- 数据成本是训练成本的10至1000倍,差距巨大
- 呼吁建立公平机制,让数据生产者获得应有回报
训练最先进的大语言模型(LLM)成本日益高昂,涉及计算、硬件、能源和工程等多重开销。然而,常被忽视且极少支付的是其训练数据背后的人力投入。每个LLM都依赖海量人类劳动:来自书籍、学术论文、代码库、社交媒体等的数万亿字精心撰写内容。本文通过研究2016至2024年间发布的64个LLM,估算若从零开始人工构建其训练数据集所需成本。即使采用极为保守的工资标准,这些数据集的成本也达到模型训练成本的10至1000倍,对大模型提供方构成重大财务负担。面对训练数据价值与未获补偿之间的巨大鸿沟,本文提出并探讨未来实现更公平实践的研究方向。
原文摘要 · Abstract (English)
Training a state-of-the-art Large Language Model (LLM) is an increasingly expensive endeavor due to growing computational, hardware, energy, and engineering demands. Yet, an often-overlooked (and seldom paid) expense is the human labor behind these models' training data. Every LLM is built on an unfathomable amount of human effort: trillions of carefully written words sourced from books, academic papers, codebases, social media, and more. This position paper aims to assign a monetary value to this labor and argues that the most expensive part of producing an LLM should be the compensation provided to training data producers for their work. To support this position, we study 64 LLMs released between 2016 and 2024, estimating what it would cost to pay people to produce their training datasets from scratch. Even under highly conservative estimates of wage rates, the costs of these models' training datasets are 10-1000 times larger than the costs to train the models themselves, representing a significant financial liability for LLM providers. In the face of the massive gap between the value of training data and the lack of compensation for its creation, we highlight and discuss research directions that could enable fairer practices in the future.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。