arXiv:2509.25045cs.CLcs.AI2025-09被引 1

用符号向量架构破解大模型内部表示,兼顾输入输出双重理解。

Hyperdimensional Probe: Decoding LLM Representations via Vector Symbolic Architectures

  • 结合符号表示与神经探针,统一输入输出分析视角。
  • 在多种模型和配置下稳定提取语义信息,揭示推理机制。
  • 适合研究大模型可解释性与内部表征的学者使用。

尽管大型语言模型(LLMs)具备强大能力,其内部表示仍不透明,现有可解释性方法或聚焦输入特征提取(如监督探针、稀疏自编码器SAEs),或关注输出分布分析(如基于logit的方法)。然而,全面理解LLM向量空间需融合两者视角,而现有方法受限于潜在特征定义。本文提出超维度探针(Hyperdimensional Probe),一种融合符号表示与神经探针的混合方法。借助向量符号架构(VSAs)与超向量代数,该方法整合了监督探针的自上而下解释性、SAE的稀疏驱动代理空间及输出导向的logit分析。通过结合传统探针的监督学习范式与SAEs的词典式表示原则,本方法实现更深层的输入特征提取,并支持输出分析。实验表明,该方法在不同模型、嵌入维度和配置下均能持续提取有意义的语义信息,在输入补全任务与问答导向文本生成两种场景中揭示概念导向的推理洞察。基于VSA的探针克服了logit分析受限于模型词汇表的缺陷,同时缓解了在有限概念空间下SAEs产生的噪声问题。该工作通过联合分析输入-输出特征,推进了对神经表示的语义理解,并统一了前人方法的互补视角。

原文摘要 · Abstract (English)

Despite their capabilities, Large Language Models (LLMs) remain opaque with limited understanding of their internal representations. Current interpretability methods either focus on input-oriented feature extraction, such as supervised probes and Sparse Autoencoders (SAEs), or on output distribution inspection, such as logit-oriented approaches. A full understanding of LLM vector spaces, however, requires integrating both perspectives, something existing approaches struggle with due to constraints on latent feature definitions. We introduce the Hyperdimensional Probe, a hybrid supervised probe that combines symbolic representations with neural probing. Leveraging Vector Symbolic Architectures (VSAs) and hypervector algebra, it unifies prior methods: the top-down interpretability of supervised probes, SAE's sparsity-driven proxy space, and output-oriented logit investigation. By combining the supervised learning paradigm of traditional probes with the dictionary-based representation principle of SAEs, our approach enables deeper input-focused feature extraction while supporting output-oriented analysis. Our experiments demonstrate that our approach consistently extracts meaningful semantic information across different LLMs, embedding sizes, and configurations, uncovering concept-oriented insights into LLM inference across two distinct scenarios: input-completion tasks and QA-focused text generation. VSA-based probing overcomes the limitations of logit-based analyses, which are constrained by the model's token vocabulary, while also mitigating the noisier interpretability outcomes often produced by SAEs in settings with a bounded conceptual feature space. By supporting a joint investigation of input-output features, this work advances the semantic understanding of neural representations while unifying the complementary perspectives of prior methods.

可解释性大模型向量符号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。