arXiv:2603.20867cs.LGcs.AI2026-03

提出语义片段新范式,破解大模型中局部语义难全局统一的难题。

Semantic Sections: An Atlas-Native Feature Ontology for Obstructed Representation Spaces

  • 用上下文图谱定义局部语义片段,支持路径一致的传播机制
  • 发现非平凡的可全局化与扭曲语义片段,打破传统向量相似性局限
  • 适合研究模型内部表示结构、解释性与跨模型语义对齐的学者

现有可解释性研究常将特征视为跨上下文共享的全局方向或坐标。我们指出,在受阻表示空间中,局部一致的语义未必能整合为全局一致特征。为此提出‘语义片段’——一种基于上下文图谱的、具备传输兼容性的局部特征族。形式化证明树支撑传播始终路径可实现,且循环一致性是真实全局化的关键标准。据此区分出树局部、可全局化与扭曲片段,其中扭曲片段捕捉了局部一致但被全纯性阻碍的语义。进一步构建发现与认证流水线:基于种子传播、重叠区同步、缺陷剪枝、循环感知分类与去重。在 Llama 3.2 3B Instruct、Qwen 2.5 3B Instruct 与 Gemma 2 2B IT 的第16层图谱中,均发现显著的语义片段群体,包括去重后仍存在的循环支持的可全局化与扭曲态。最重要的是,原始全局向量相似性无法恢复语义身份:即使已认证的可全局化片段也表现出低跨图谱符号余弦相似度,而基线方法仅能恢复少量真实同段对,常在中等阈值处坍塌。相比之下,基于片段的身份恢复在认证支持上达到完美效果。这些结果支持语义片段作为受阻场景下更优的特征本体。

原文摘要 · Abstract (English)

Recent interpretability work often treats a feature as a single global direction, dictionary atom, or latent coordinate shared across contexts. We argue that this ontology can fail in obstructed representation spaces, where locally coherent meanings need not assemble into one globally consistent feature. We introduce an atlas-native replacement object, the semantic section: a transport-compatible family of local feature representatives defined over a context atlas. We formalize semantic sections, prove that tree-supported propagation is always pathwise realizable, and show that cycle consistency is the key criterion for genuine globalization. This yields a distinction between tree-local, globalizable, and twisted sections, with twisted sections capturing locally coherent but holonomy-obstructed meanings. We then develop a discovery-and-certification pipeline based on seeded propagation, synchronization across overlaps, defect-based pruning, cycle-aware taxonomy, and deduplication. Across layer-16 atlases for Llama 3.2 3B Instruct, Qwen 2.5 3B Instruct, and Gemma 2 2B IT, we find nontrivial populations of semantic sections, including cycle-supported globalizable and twisted regimes after deduplication. Most importantly, semantic identity is not recovered by raw global-vector similarity. Even certified globalizable sections show low cross-chart signed cosine similarity, and raw similarity baselines recover only a small fraction of true within-section pairs, often collapsing at moderate thresholds. By contrast, section-based identity recovery is perfect on certified supports. These results support semantic sections as a better feature ontology in obstructed regimes.

可解释性表示学习大模型语义结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。