大模型能从乱码英文中还原原意,揭示语言结构的深层约束。
The Astonishing Ability of Large Language Models to Parse Jabberwockified Language
- 用无意义词替换关键词,模型仍能还原句子语义。
- 在严重退化文本下,翻译准确率接近原句,体现超强泛化能力。
- 适合研究语言认知、模型推理机制与自然语言理解的学者。
我们发现大型语言模型(LLMs)具备惊人的能力,能够从严重退化的英语文本中恢复语义。当内容词被随机替换为无意义字符串时,例如“At the ghybe of the swuint, we are haiveed to Wourge Phrear-gwurr, who sproles into an ghitch flount with his crurp”,模型可将其翻译为接近原文的正常英语,如“At the start of the story, we meet a man, Chow, who moves into an apartment building with his wife.” 这表明,结构线索(如形态句法、封闭类词)对词汇意义的约束远超预期。尽管模型在解析“Jabberwockified”英语方面表现远超人类,但其结果对理解语言结构具有重要意义,并暗示高效的语言处理——无论是生物还是人工系统——可能依赖于句法、词汇语义与常识知识的紧密整合。
原文摘要 · Abstract (English)
We show that large language models (LLMs) have an astonishing ability to recover meaning from severely degraded English texts. Texts in which content words have been randomly substituted by nonsense strings, e.g., "At the ghybe of the swuint, we are haiveed to Wourge Phrear-gwurr, who sproles into an ghitch flount with his crurp", can be translated to conventional English that is, in many cases, close to the original text, e.g., "At the start of the story, we meet a man, Chow, who moves into an apartment building with his wife." These results show that structural cues (e.g., morphosyntax, closed-class words) constrain lexical meaning to a much larger degree than imagined. Although the abilities of LLMs to make sense of "Jabberwockified" English are clearly superhuman, they are highly relevant to understanding linguistic structure and suggest that efficient language processing either in biological or artificial systems likely benefits from very tight integration between syntax, lexical semantics, and general world knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。