arXiv:2506.18011cs.LGcs.AI2025-06被引 1

通过微小词元扰动,揭示Transformer模型中信息传播机制。

Probing the Embedding Space of Transformers via Minimal Token Perturbations

  • 用极小词元扰动探测嵌入空间变化。
  • 稀有词元扰动导致更大嵌入偏移。
  • 深层网络中信息逐渐混合,适合模型解释研究。

理解信息在Transformer模型中的传播路径是可解释性研究的关键挑战。本文通过分析极小词元扰动对嵌入空间的影响,发现稀有词元通常引发更大的嵌入偏移。同时研究扰动在各层间的传播,表明输入信息在深层网络中逐渐相互混合。结果验证了将模型前几层作为解释代理的常见假设。本工作提出将词元扰动与嵌入空间偏移结合,作为一种强有力的模型可解释性工具。

原文摘要 · Abstract (English)

Understanding how information propagates through Transformer models is a key challenge for interpretability. In this work, we study the effects of minimal token perturbations on the embedding space. In our experiments, we analyze the frequency of which tokens yield to minimal shifts, highlighting that rare tokens usually lead to larger shifts. Moreover, we study how perturbations propagate across layers, demonstrating that input information is increasingly intermixed in deeper layers. Our findings validate the common assumption that the first layers of a model can be used as proxies for model explanations. Overall, this work introduces the combination of token perturbations and shifts on the embedding space as a powerful tool for model interpretability.

模型可解释性Transformer嵌入空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。