arXiv:2510.19131cs.CLcs.LG2025-10被引 2

通过频谱分析揭示模型在语音切换时的计算指纹,发现架构差异可被量化检测。

Training-Free Spectral Fingerprints of Voice Processing in Transformers

  • 用注意力图谱的代数连通性变化追踪早期层的计算响应模式。
  • Phi-3-Mini在英语中出现显著连通性下降(Δλ₂≈-0.446),其他语言影响小。
  • 无需训练即可诊断模型偏见,适用于评估模型可靠性与推理模式差异。

不同Transformer架构通过各异的连接模式实现相同的语言计算,形成可被频谱分析捕捉的“计算指纹”。通过对注意力诱导的标记图进行图信号处理,我们在20种语言和三个模型族中,于预设早期窗口(第2–5层)追踪语音切换下的代数连通性变化(弗里德曼值,Δλ₂)。结果揭示清晰的架构特征:Phi-3-Mini在英语中表现出显著的早期层破坏(平均Δλ₂≈-0.446),而其他19种语言影响极小,符合其主要面向英语的公开文档;Qwen2.5-7B呈现微小且分布式的位移,以形态丰富的语言为最大;LLaMA-3.2-1B则显示系统但微弱的响应。这些频谱特征与行为差异高度相关(Phi-3: r = -0.976),并受特定注意力头消融调节,表明其与早期注意力结构的功能关联。研究支持训练侧重会留下可检测的计算印记,即语法变换中的可测量连接模式。该框架还可区分推理模式,具有无需训练、简单有效的诊断潜力,适用于揭示架构偏差与支持模型可靠性分析。

原文摘要 · Abstract (English)

Different transformer architectures implement identical linguistic computations via distinct connectivity patterns, yielding model imprinted ``computational fingerprints'' detectable through spectral analysis. Using graph signal processing on attention induced token graphs, we track changes in algebraic connectivity (Fiedler value, $Δλ_2$) under voice alternation across 20 languages and three model families, with a prespecified early window (layers 2--5). Our analysis uncovers clear architectural signatures: Phi-3-Mini shows a dramatic English specific early layer disruption ($\overline{Δλ_2}_{[2,5]}\!\approx\!-0.446$) while effects in 19 other languages are minimal, consistent with public documentation that positions the model primarily for English use. Qwen2.5-7B displays small, distributed shifts that are largest for morphologically rich languages, and LLaMA-3.2-1B exhibits systematic but muted responses. These spectral signatures correlate strongly with behavioral differences (Phi-3: $r=-0.976$) and are modulated by targeted attention head ablations, linking the effect to early attention structure and confirming functional relevance. Taken together, the findings are consistent with the view that training emphasis can leave detectable computational imprints: specialized processing strategies that manifest as measurable connectivity patterns during syntactic transformations. Beyond voice alternation, the framework differentiates reasoning modes, indicating utility as a simple, training free diagnostic for revealing architectural biases and supporting model reliability analysis.

Transformer频谱分析模型诊断语音处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。