翻译劳动被当作AI训练数据,却未获认可与回报。
Translators as Invisible Teachers of AI: Copyright, Translation Memory, and the Political Economy of Linguistic Data
- 翻译记忆库成AI训练核心数据,但译者无权主张权益。
- 法律允许仅提取数据特征而不需使用原作,形成隐性剥削。
- 适合关注AI伦理、数据权利与译者权益的研究者阅读。
本文探讨译者劳动如何转化为人工智能时代的基础数据资本。翻译记忆库(TM)与平行语料库保留源语文本与目标语文本的一一对应关系,因此成为机器翻译极为宝贵的监督训练数据。统计机器翻译(SMT)、神经机器翻译(NMT)、Transformer架构及多语言大模型(LLMs)的发展均离不开此类翻译数据的积累。然而,译者的成果被作为合同交付物购买,被拆解为技术对象,并以版权法下的“信息分析”数据形式处理,导致其道德、创作与经济归属被剥夺。论文提出两个概念:一是‘不消费的占有’——作品未被阅读或观看,仅用于提取统计特征,该行为在日本著作权法第30-4条下合法化;二是‘译者隐形教师化’——译者通过构建翻译记忆库、后期编辑与质量评估,无形中成为AI的教师却未获承认。基于从译者经语言服务提供商(LSPs)到平台与模型开发者的数据供应链,对比日本、欧洲与美国法律框架,区分开放与专有模型,结合模型坍塌时代人工数据的溢价地位,论文追问译者真正担忧为何,并指向再分配设计的具体方向。
原文摘要 · Abstract (English)
This paper examines how the labour of translators has been transformed into foundational data capital for the age of artificial intelligence (AI). Translation memories (TM) and parallel corpora preserve a one-to-one correspondence between source and target text and therefore constitute extraordinarily valuable supervised training data for machine translation. The development of statistical machine translation (SMT), neural machine translation (NMT), the Transformer architecture, and multilingual large language models (LLMs) cannot be disentangled from the accumulation of such translation data. And yet, translators' renditions have been bought as deliverables under contract, segmented as technical objects, and processed as "information analysis" data under copyright law -- losing their moral, creative, and economic attribution to the translators who produced them. The paper develops two concepts to capture this process. The first is appropriation without consumption: a mode of use in which works are not read, viewed, or listened to, but only mined for statistical features -- a use that is legitimated under Article 30-4 of the Japanese Copyright Act. The second is the invisible teacherisation of translators: the process by which translators, through the construction of translation memories, post-editing, and quality assessment, have functioned as teachers of AI without recognition as such. Drawing on the data supply chain that runs from translators through language service providers (LSPs) and platforms to model developers, on a comparative reading of Japanese, European, and United States legal frameworks, on the distinction between open and proprietary AI models, and on the premium status that human-generated data has acquired in the era of model collapse, the paper asks what translators are actually afraid of, and points toward concrete directions for redistributive design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。