arXiv:2603.26207cs.CLcs.AI2026-03被引 2

探讨大模型语义如何体现整体性,回应稀疏自编码器带来的挑战

Sparse Auto-Encoders and Holism about Large Language Models

  • 从分布语义出发,论证大模型隐含语义整体性
  • 指出稀疏自编码器发现可解释特征,暗示语义可分解
  • 强调只要特征可数,整体性观点仍成立,适合哲学与认知科学读者

大语言模型(LLM)技术是否暗示了一种元语义图景——即词语与复杂表达如何获得意义?一种温和的方法是通过考察LLM捕捉语言意义时所依赖的假设,来评估其合理性(Grindrod, 2026a, 2026b)。此前已有观点认为,由于采用分布语义,LLM体现了一种意义整体性(Grindrod, 2023;Grindrod et al., forthcoming)。然而,最近的机制可解释性研究对此提出挑战:稀疏自编码器在高维空间中发现了大量可解释的潜在特征,这似乎支持语义可分解的图景。本文首先回顾支持整体性的原始理由(第1节),再介绍稀疏自编码器生成特征的最新工作,并说明这些特征如何暗示分解式意义观(第2节)。接着深入分析此类特征的本质(第3节),最后重申格里德罗德等人捍卫的整体性图景依然成立,前提是这些特征是可数的(第4节)。

原文摘要 · Abstract (English)

Does Large Language Model (LLM) technology suggest a meta-semantic picture i.e. a picture of how words and complex expressions come to have the meaning that they do? One modest approach explores the assumptions that seem to be built into how LLMs capture the meanings of linguistic expressions as a way of considering their plausibility (Grindrod, 2026a, 2026b). It has previously been argued that LLMs, in employing a form of distributional semantics, adopt a form of holism about meaning (Grindrod, 2023; Grindrod et al., forthcoming). However, recent work in mechanistic interpretability presents a challenge to these arguments. Specifically, the discovery of a vast array of interpretable latent features within the high dimensional spaces used by LLMs potentially challenges the holistic interpretation. In this paper, I will present the original reasons for thinking that LLMs embody a form of holism (section 1), before introducing recent work on features generated through sparse auto-encoders, and explaining how the discovery of such features suggests an alternative decompositional picture of meaning (section 2). I will then respond to this challenge by considering in greater detail the nature of such features (section 3). Finally, I will return to the holistic picture defended by Grindrod et al. and argue that the picture still stands provided that the features are countable (section 4).

大模型语义哲学可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。