arXiv:2604.27169cs.CLcs.LG2026-04

大模型的语义特征空间与人类心理关联高度一致,揭示了语义的几何结构。

Semantic Structure of Feature Space in Large Language Models

论文配图:Semantic Structure of Feature Space in Large Language Models
图 1 · 摘自论文原文
  • 用32个语义轴投影360个词,发现模型特征与人类评分高度相关。
  • 语义轴间的余弦相似性可预测人类对语义尺度的相关性。
  • 语义变异集中在低维子空间,符合人类语义关联规律,适合理解模型认知机制。

我们发现大语言模型隐藏状态中语义特征的几何关系与人类心理关联高度一致。构建了360个词对应的特征向量,并将其投影到32个语义轴(如美-丑、软-硬)上,发现这些投影与人类对词汇在相应语义尺度上的评分高度相关。其次,语义轴之间的余弦相似性能有效预测这些语义尺度在调查中的相关性。第三,32个语义轴的大部分方差集中在一个低维子空间,重现了人类语义关联的典型模式。最后,沿某一语义轴操控一个词,会引发其在其他语义尺度上的评分变化,且变化幅度与对应语义轴的余弦相似性成正比。这些结果表明,语义特征应通过其几何关系和所形成的有意义子空间来理解。

原文摘要 · Abstract (English)

We show that the geometric relations between semantic features in large language models' hidden states closely mirror human psychological associations. We construct feature vectors corresponding to 360 words and project them on 32 semantic axes (e.g. beautiful-ugly, soft-hard), and find that these projections correlate highly with human ratings of those words on the respective semantic scales. Second, we find that the cosine similarities between the semantic axes themselves are highly predictive of the correlations between these scales in the survey. Third, we show that substantial variance across the 32 semantic axes lies on a low-dimensional subspace, reproducing patterns typical of human semantic associations. Finally, we demonstrate that steering a word on one semantic axis causes spillover effects on the model's rating of that word on other semantic scales proportionate to the cosine similarity between those semantic axes. These findings suggest that features should be understood not only in isolation but through their geometric relations and the meaningful subspaces they form.

语义空间大模型几何结构心理关联

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。