arXiv:2609.03577cs.CL2026-09

反思语言模型究竟是技术产品还是语言研究工具。

Language, Language Models, and What We're Talking About

论文配图:Language, Language Models, and What We're Talking About
图 1 · 摘自论文原文
  • 区分语言模型作为技术产品与语言研究工具的定位差异。
  • 指出用翻译/合成数据训练的模型未必真正代表目标语言。
  • 适合关注NLP本质与语言建模哲学的研究者阅读。

语言模型常被当作技术产物讨论,但其本质受训练数据所传达的语言世界影响。以意大利语模型为例,本文关注在翻译和合成数据上训练并持续优化后形成的系统,以及在同样不自然的数据上进行测试的意义。这些模型究竟是意大利语模型吗?是语言模型吗?自然语言处理是否仍关心语言本身?这些问题引出更实际的问题:我们真正希望语言模型生成什么样的语言?作者认为,若不先厘清语言模型作为技术产品与语言研究工具的界限,就无法回答此问题。答案可能多样,我们谈论的语言也可能多样,局面未必如想象般悲观。

原文摘要 · Abstract (English)

Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result from training and specialising models on translated and synthetic data, and further curating them, and to the meaning of testing them on equally unnatural data. Are these eventually models of Italian? Are they models of language? Does NLP still care about language? These questions yield another, more concrete question: what language do we actually want language models to produce? I argue that this question cannot be answered if we do not first consider a clearer distinction between language models designed as technical products and language models designed as tools for studying language itself. The answers then might be diverse, the languages we are talking about might be diverse, and the picture might not be as pessimistic as we fear.

语言模型NLP哲学语言本质

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。