用拓扑方法揭示大模型对抗攻击如何压缩内部表征空间
The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
- 用持久同调分析模型内部表征的几何拓扑变化
- 对抗输入使隐空间结构简化,从多小尺度特征变为少数大尺度主导
- 该现象在不同攻击和模型中一致,适合研究模型安全与可解释性
现有大语言模型可解释性方法多关注线性方向或孤立特征,忽略了模型表示的高维、关系性和非线性几何。本文采用持久同调(PH)刻画对抗输入如何重塑大模型内部表示空间的几何与拓扑结构。针对六种模型(参数量3.8B至70B)在两种攻击模式(间接提示注入与后门微调)下的表现,发现一种稳定的拓扑特征始终存在:对抗输入引发拓扑压缩,即隐空间结构趋于简单,将原本多样、紧凑的小尺度特征坍缩为少数主导的大尺度特征。该特征具有架构无关性,在网络早期出现,且在各层间具有高度区分性。通过量化激活点云形状与神经元级信息流动,本框架揭示了表征变化的几何不变量,补充了现有线性可解释方法的不足。
原文摘要 · Abstract (English)
Existing interpretability methods for Large Language Models (LLMs) predominantly capture linear directions or isolated features. This overlooks the high-dimensional, relational, and nonlinear geometry of model representations. We apply persistent homology (PH) to characterize how adversarial inputs reshape the geometry and topology of internal representation spaces of LLMs. This phenomenon, especially when considered across operationally different attack modes, remains poorly understood. We analyze six models (3.8B to 70B parameters) under two distinct attacks, indirect prompt injection and backdoor fine--tuning, and show that a consistent topological signature persists throughout. Adversarial inputs induce topological compression, where the latent space becomes structurally simpler, collapsing the latent space from varied, compact, small-scale features into fewer, dominant, large-scale ones. This signature is architecture-agnostic, emerges early in the network, and is highly discriminative across layers. By quantifying the shape of activation point clouds and neuron-level information flow, our framework reveals geometric invariants of representational change that complement existing linear interpretability methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。