发现语言模型采样误差与内部状态的几何关系
A geometric relation of the error introduced by sampling a language model's output distribution to its internal state

- 用微分几何建模令牌空间,提取由嵌入结构决定的1形式
- 曲率在国际象棋任务中与世界模型耦合,体现棋局区域与棋子重要性
- 揭示了模型内部表征可直接从令牌空间几何读取
GPT类语言模型在生成时对单个标记变化敏感,尤其当预测分布分散于多个标记时。将这种敏感性视为几何属性,我们推导出一个仅依赖令牌嵌入几何的$\/mathfrak{so}(n)$-值1形式。尽管其起源纯粹几何,但其曲率具有语义意义:在国际象棋推理任务中,曲率与现成指令微调模型的世界模型耦合,变换按棋盘区域聚类并尊重棋子重要性。研究结果表明,令牌空间几何直接反映了模型对问题的内部表征。
原文摘要 · Abstract (English)
GPT-style language models are sensitive to single-token changes at generation points where the predicted probability distribution is spread across multiple tokens. Viewing this sensitivity as a geometric property, we derive an $\mathfrak{so}(n)$-valued 1-form that depends only on the geometry of the token embeddings. Despite this purely geometric origin, we show that its curvature is semantically meaningful: On chess reasoning tasks, the curvature couples to the world model of an off-the-shelf instruction-tuned model, with transformations clustering by board region and respecting piece importance. Our findings suggest that token space geometry directly reflects how models internally represent problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。