用神经混沌方法捕捉冻结Transformer输入空间的响应指纹。
ChaosProbe: A Neurochaotic Lens on Frozen Transformer Input-Embedding Spaces

- 设计混沌轨迹变换,生成输入嵌入的响应签名。
- 四模型在80个提示下,相关性指标均识别出同家族最近邻对。
- 适合研究模型内部表征结构与稳定性分析的读者。
Transformer模型通常通过其性能或下游任务表现来理解,但冻结的输入嵌入空间也可通过可控确定性探测器在上下文计算前的响应进行分析。基于此响应视角,我们提出 extit{ChaosProbe},一种受神经混沌启发的确定性方法,用于构建冻结Transformer输入嵌入空间的响应指纹。对每个提示级嵌入矩阵,ChaosProbe应用基于混沌轨迹的变换,并以发放率和熵通道响应为互补表示度量,生成固定长度签名。在包含80个中性提示和四个预训练模型(GPT-2、DistilGPT2、BERT-base-uncased、RoBERTa-base)的小规模验证中,皮尔逊相关、斯皮尔曼相关和余弦相似性均恢复全部四组同家族最近邻配对及两组预期互属家族对;欧氏距离恢复三组最近邻和一组互属对。配对自举重采样支持皮尔逊与斯皮尔曼配对在观察提示集上的稳定性,签名有效性检验表明恒定或坍缩响应未主导报告指纹。结果提供了一种群体依赖的初步证据:确定性神经混沌响应签名可揭示冻结Transformer输入嵌入空间间的广泛结构。
原文摘要 · Abstract (English)
Transformer models are most often understood through what they do: their benchmark performance, generation quality, or behavior on downstream tasks. Yet frozen transformer input-embedding spaces may also be examined through their responses to a controlled deterministic probe before contextual computation or task-specific adaptation. Guided by this response-based view, we introduce \emph{ChaosProbe}, a deterministic neurochaos-inspired method for constructing response-based fingerprints of frozen transformer input-embedding spaces. For each prompt-level embedding matrix, ChaosProbe applies a chaotic trajectory-based transformation and summarizes its Firing Rate and Entropy channel responses with complementary representation-level measures, producing a fixed-length signature for each model. In a bounded proof-of-concept study of $80$ neutral prompts and four pretrained models---GPT-2, DistilGPT2, BERT-base-uncased, and RoBERTa-base---Pearson correlation, Spearman correlation, and cosine similarity each recover all four same-family nearest-neighbor assignments and both expected mutual family pairs. Euclidean distance recovers three of the four assignments and one of the two mutual family pairs. Paired bootstrap resampling supports the stability of the Pearson and Spearman pairings over the observed prompt set, and signature-validity checks show that constant or collapsed responses do not dominate the reported fingerprints. These results provide a cohort-dependent proof of concept that deterministic neurochaotic response signatures can expose broad structure among frozen transformer input-embedding spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。