arXiv:2508.14067cs.CLcs.LG2025-08被引 5

研究大模型如何处理标点和逻辑推理,发现不同模型差异显著。

Punctuation and Predicates in Language Models

  • 通过干预实验分析标点在各层的作用机制
  • GPT-2中标点对性能至关重要,其他模型则不然
  • 条件句与全称命题被模型以不同方式处理

本文探究大语言模型(LLMs)中信息的收集位置与传播路径。首先分析标点符号的计算重要性,此前研究认为其为注意力聚集点与记忆辅助。利用干预技术评估标点在GPT-2、DeepSeek和Gemma各层中的必要性与充分性。结果显示:在GPT-2中,标点在多个层中既必要又充分;而在DeepSeek中作用较弱,在Gemma中几乎无影响。进一步考察模型是否对输入成分(如主语、形容词、标点、完整句子)形成早期静态摘要并复用,或持续敏感于其变化。通过互换干预与层交换实验发现,条件句(if, then)与全称量词(for all)的处理方式截然不同。研究揭示了标点与推理在模型内部的运作机制,对可解释性具有重要意义。

原文摘要 · Abstract (English)

In this paper we explore where information is collected and how it is propagated throughout layers in large language models (LLMs). We begin by examining the surprising computational importance of punctuation tokens which previous work has identified as attention sinks and memory aids. Using intervention-based techniques, we evaluate the necessity and sufficiency (for preserving model performance) of punctuation tokens across layers in GPT-2, DeepSeek, and Gemma. Our results show stark model-specific differences: for GPT-2, punctuation is both necessary and sufficient in multiple layers, while this holds far less in DeepSeek and not at all in Gemma. Extending beyond punctuation, we ask whether LLMs process different components of input (e.g., subjects, adjectives, punctuation, full sentences) by forming early static summaries reused across the network, or if the model remains sensitive to changes in these components across layers. Extending beyond punctuation, we investigate whether different reasoning rules are processed differently by LLMs. In particular, through interchange intervention and layer-swapping experiments, we find that conditional statements (if, then), and universal quantification (for all) are processed very differently. Our findings offer new insight into the internal mechanisms of punctuation usage and reasoning in LLMs and have implications for interpretability.

语言模型标点分析推理机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。