arXiv:2606.02750cs.CL2026-06

发现大模型表征受词汇重叠影响深远,削弱语义理解能力。

On the Persistent Effects of Lexicality in Large Language Models

  • 通过对抗性测试量化词汇重叠对表征的影响
  • 发现词汇效应贯穿模型深层,且在中层同时削弱词形与语义信号
  • 揭示该现象影响摘要生成与模型编辑等下游任务

从大语言模型(LLMs)中提取的表征在众多下游应用中起关键作用,但其结构常受词汇重叠而非语义内容影响。我们研究了这种词汇影响与语义内容的关系及其对下游任务的含义。通过多种对抗性语义压力测试,并结合信息论视角,发现词汇影响贯穿模型深度,且在不同架构、训练方式和目标函数下均一致存在,包括专为语义相似性训练的模型。此外,在模型中层观察到词汇与语义信号同时退化,表明此区域为表征质量低下的过渡阶段。通过摘要生成和模型编辑案例,进一步验证了词汇影响对下游应用的实际危害。

原文摘要 · Abstract (English)

Representations extracted from large language models (LLMs) play an important role in many downstream applications. However, the structure of these representations is often influenced by lexical overlap rather than semantic content. Our understanding of the relationship between this lexical influence and semantic content, and its implications for downstream tasks, remains limited. In this work, we investigate representations to quantify the effect of lexical overlap relative to semantic content. We consider several adversarial semantic stress tests and further connect our findings to the information theory perspective. We find that lexical influence extends across the depth of models, consistently across architectures, training regimes, and objective functions, including the models trained for semantic similarity. Moreover, we observe a mid-depth region in which both lexical and semantic signals degrade simultaneously, indicating a transitional regime where representations are poor for both surface form and meaning. We further demonstrate the effect of lexical influence on downstream uses of LLMs using summarization and model editing as a case study.

大模型表征词汇重叠语义退化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。