通过几何轨迹分析大模型如何响应语义关注变化。
Curved Inference: Concern-Sensitive Geometry in Large Language Model Residual Streams
- 用残差流曲率与显著性度量语义关注的几何变化。
- LLaMA 模型在关注强度增强时曲率与显著性显著上升。
- 适合研究模型对齐、抽象能力与推理动态的学者参考。
我们提出 Curved Inference——一种几何可解释性框架,用于追踪大语言模型残差流轨迹在语义关注变化下的弯曲情况。在涵盖情感、道德、视角、逻辑、身份、环境及无意义等领域的20组匹配提示下,我们分析了 Gemma3-1b 与 LLaMA3.2-3b 模型,采用五种原空间度量,重点考察曲率(κ_i)和显著性(S(t))。这些度量基于未嵌入矩阵导出的拉回语义度量计算,确保所有测量反映对齐词元的几何结构而非原始坐标。结果表明,关注变化会可靠改变两个模型的内部激活轨迹;其中 LLaMA 在关注强度增加时表现出一致且统计显著的曲率与显著性增长;而 Gemma 虽然也响应关注,但在中等与强关注间区分较弱。研究支持大模型几何的两层观:嵌入空间中的潜在概念结构,以及由提示特定推理塑造的上下文轨迹。Curved Inference 揭示了模型在深度中如何导航、重定向或强化语义意义,为诊断对齐、抽象与涌现推理动态提供了原则性方法,为理解语义抽象与模型对齐提供了新视角。
原文摘要 · Abstract (English)
We propose Curved Inference - a geometric Interpretability framework that tracks how the residual stream trajectory of a large language model bends in response to shifts in semantic concern. Across 20 matched prompts spanning emotional, moral, perspective, logical, identity, environmental, and nonsense domains, we analyse Gemma3-1b and LLaMA3.2-3b using five native-space metrics, with a primary focus on curvature (\k{appa}_i) and salience (S(t)). These metrics are computed under a pullback semantic metric derived from the unembedding matrix, ensuring that all measurements reflect token-aligned geometry rather than raw coordinate structure. We find that concern-shifted prompts reliably alter internal activation trajectories in both models - with LLaMA exhibiting consistent, statistically significant scaling in both curvature and salience as concern intensity increases. Gemma also responds to concern but shows weaker differentiation between moderate and strong variants. Our results support a two-layer view of LLM geometry - a latent conceptual structure encoded in the embedding space, and a contextual trajectory shaped by prompt-specific inference. Curved Inference reveals how models navigate, reorient, or reinforce semantic meaning over depth, offering a principled method for diagnosing alignment, abstraction, and emergent inference dynamics. These findings offer fresh insight into semantic abstraction and model alignment through the lens of Curved Inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。