arXiv:2508.10003cs.CLcs.AI2025-08被引 9

发现大模型词嵌入中语义结构呈低维特征,与人类认知高度一致。

Semantic Structure in Large Language Model Embeddings

  • 用反义对定义语义方向,投影可匹配人类评分
  • 语义特征压缩至3维子空间,信息损失极小
  • 语义扰动会通过余弦相似性引发连锁效应

心理研究表明,人类对词汇在各类语义维度上的评分可被降维为低维表示且信息损失极小。我们发现大语言模型(LLMs)的嵌入矩阵中编码的语义关联也具有类似结构。通过反义对(如 kind - cruel)定义的语义方向,词的投影与人类评分高度相关,并进一步表明这些投影可有效压缩至3维子空间,与人类问卷数据导出的模式高度吻合。此外,沿某一语义方向移动标记符时,会对几何对齐特征产生非目标影响,其强度与特征间的余弦相似性成正比。这表明,语义特征在大模型中的纠缠方式与人类语言认知相似,大量看似复杂的语义信息实则具有惊人低维性。考虑此结构对避免特征调控中的意外后果至关重要。

原文摘要 · Abstract (English)

Psychological research consistently finds that human ratings of words across diverse semantic scales can be reduced to a low-dimensional form with relatively little information loss. We find that the semantic associations encoded in the embedding matrices of large language models (LLMs) exhibit a similar structure. We show that the projections of words on semantic directions defined by antonym pairs (e.g. kind - cruel) correlate highly with human ratings, and further find that these projections effectively reduce to a 3-dimensional subspace within LLM embeddings, closely resembling the patterns derived from human survey responses. Moreover, we find that shifting tokens along one semantic direction causes off-target effects on geometrically aligned features proportional to their cosine similarity. These findings suggest that semantic features are entangled within LLMs similarly to how they are interconnected in human language, and a great deal of semantic information, despite its apparent complexity, is surprisingly low-dimensional. Furthermore, accounting for this semantic structure may prove essential for avoiding unintended consequences when steering features.

语义结构嵌入分析低维表征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。