arXiv:2602.09784cs.LGcs.CL2026-02被引 1

答案词元编码生成路径方向,揭示了模型可解释与可控性的几何统一

Circuit Fingerprints: How Answer Tokens Encode Their Geometrical Path

  • 通过几何对齐发现电路结构,无需梯度或因果干预
  • 在4个模型族上达到与梯度方法相当的电路发现效果
  • 同一方向可实现精准情感控制,准确率达69.8%

Transformer中的电路发现与激活操控曾作为独立研究方向,但二者均作用于同一表征空间。我们证明它们遵循同一几何规律:答案词元在孤立处理时,编码了生成自身的方向。这一电路指纹假说使无需梯度或因果干预即可实现电路发现——仅靠几何对齐便能恢复与梯度方法相当的结构。我们在标准基准(IOI、SVA、MCQA)上验证该方法,涵盖四个模型家族,表现相当。相同的方向不仅能识别电路组件,还能实现可控操控:情感分类准确率达69.8%,显著优于指令提示的53.1%,且保持事实准确性。这表明,变压器电路本质上是几何结构,可解释性与可控性实为同一对象的两个方面。

原文摘要 · Abstract (English)

Circuit discovery and activation steering in transformers have developed as separate research threads, yet both operate on the same representational space. Are they two views of the same underlying structure? We show they follow a single geometric principle: answer tokens, processed in isolation, encode the directions that would produce them. This Circuit Fingerprint hypothesis enables circuit discovery without gradients or causal intervention -- recovering comparable structure to gradient-based methods through geometric alignment alone. We validate this on standard benchmarks (IOI, SVA, MCQA) across four model families, achieving circuit discovery performance comparable to gradient-based methods. The same directions that identify circuit components also enable controlled steering -- achieving 69.8\% emotion classification accuracy versus 53.1\% for instruction prompting while preserving factual accuracy. Beyond method development, this read-write duality reveals that transformer circuits are fundamentally geometric structures: interpretability and controllability are two facets of the same object.

可解释性几何结构可控性Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。