用代数几何方法解决大模型词嵌入的语义歧义问题
TokenBlowUp: Resolving Representational Singularities in LLM Token Spaces via Monoidal Transformations
- 基于格理论对词嵌入奇异点进行代数爆破处理
- 证明新空间几何结构可消除原始语义不稳定性
- 适合研究模型内在表征与几何学习的学者
近期研究表明,大型语言模型的词嵌入空间违背了基础流形假设,存在多义词附近的几何奇异点,导致表征不稳定。现有方法依赖平滑数据流形假设,无法应对此类结构性缺陷。本文以格理论形式化该问题,提出在每个奇异点处应用方案论爆破操作。该过程将奇异点替换为例外除子,我们将其识别为方向投影空间——一个容纳词义消歧的规范几何空间。此“表征去奇异化”构建了新的嵌入几何景观,并证明了新空间的几何正则性,确保原始病态问题被解决。最后,我们讨论了该框架的架构启示,主张从静态查表转向动态、几何驱动的计算范式。
原文摘要 · Abstract (English)
Recent work has provided compelling evidence challenging the foundational manifold hypothesis for the token embedding spaces of Large Language Models (LLMs). These findings reveal the presence of geometric singularities around polysemous tokens, which can lead to representational instability. Existing methodologies, which presuppose a smooth data manifold, are ill-equipped to address such intrinsic structural flaws. In this paper, we formalize this problem in the language of scheme theory and propose a rigorous resolution by applying the scheme-theoretic blow-up at each singular point. This procedure replaces a singular point in the ambient affine scheme with its exceptional divisor, which we identify as a canonical geometric space -- a projective space of directions -- that houses the disambiguated semantic meanings of the token. This process of ``representational desingularization'' constructs a new geometric landscape for embeddings. We prove a formal theorem guaranteeing the geometric regularization of this new space, showing that the original pathologies are resolved. Finally, we outline the architectural implications of our framework, arguing for a paradigm shift from static look-ups to dynamic, geometrically-grounded computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。