多语言大模型正导致语言多样性萎缩,需警惕其隐性危害。
Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs
- 通过自循环训练数据引发模型坍塌,压缩语言表达
- 低概率语言特征逐渐消失,文化内涵随之流失
- 呼吁重视语言多样性,重塑NLP研究方向
多语言大语言模型正深刻改变技术对语言的影响方式。以往技术仅辅助人类,如今却使写作任务被直接外包给模型,导致语言被更直接地重塑。尽管它们提供快速信息获取和流畅输出,但其背后潜藏一种微妙而危险的趋势:语言多样性的逐步衰退与丧失。本文探讨了模型坍塌现象——在翻译技术等场景中,自动生成的数据反复回流训练集,形成自我强化的训练循环,致使数据分布扭曲,低概率语言现象(如特定语法结构、文化表达)被边缘化。借鉴计算机视觉、自然语言处理与机器翻译领域的最新研究,本文指出我们语言分布中的“长尾”特征正在消失,连同其所承载的故事与身份认同一并消逝。这是一篇呼吁抵抗语言同质化的立场论文,倡导将自然语言处理重新定位为鼓励、珍视并保护多元语言表达与创造力的领域。
原文摘要 · Abstract (English)
Multilingual Large Language Models considerably changed how technologies influence language. While previous technologies could mediate or assist humans, there is now a tendency to offload the task of writing itself to these technologies, enabling models to change our languages more directly. While they provide us quick access to information and impressively fluent output, beneath their (apparent) sophistication lies a subtle, insidious threat: the gradual decline and loss of linguistic diversity. In this position paper, I explore how model collapse, with a particular focus on translation technology, can lead to the loss of linguistic forms, grammatical features, and cultural nuance. Model collapse refers to the consequences of self-consuming training loops, where automatically generated data (re-)enters the training data, leading to a gradual distortion of the data distribution and the underrepresentation of low-probability linguistic phenomena. Drawing on recent work in Computer Vision, Natural Language Processing and Machine Translation, I argue that the many tails of our linguistic distributions might be vanishing, and with them, the narratives and identities they carry. This paper is a call to resist linguistic flattening and to reimagine Natural Language Processing as a field that encourages, values and protects expressive multilingual diversity and creativity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。