通过设计测量方法,揭示模型内部机制与干预效果。
Mechanistic Tomography: Designed Measurement for Control-Oriented Interpretability

- 构建统一测量框架,用线性模型还原内部机制与干预响应。
- 在双隐马尔可夫模型中,观测误差直接导致控制误差上升。
- 适用于需要精准干预解释的场景,如大模型安全对齐研究。
机制可解释性旨在获取模型未显式暴露的内部量:表示状态、组件效应、交互关系及干预响应。补丁法、梯度、海塞向量积与子集干预在不同访问假设下提供不同测量,针对不同目标。本文将其共性结构形式化为机制断层扫描:面向控制的可解释性设计测量。对于选定基底与干预族,测量形式为 y = Ax + w,其中 A 表示干预设计,x 为目标映射,w 包含非线性响应、采样误差与基底错配。该语言给出实用流程:从成本最低测量开始,于保留干预上测试目标尺度,校准简单偏差,当结构残差仍存在时扩展测量族。控制作为严苛验证环境,因估计值指导干预即充当观察者。在双隐马尔可夫模型中,控制误差随观察误差上升,而目标改善可能掩盖干扰态变化。仅前向访问下,稀疏聚合测量以更少干预恢复有限效应图,优于坐标补丁;具梯度访问时,有限探针提升局部归因图。提升测量与海塞向量积能捕捉一阶图遗漏的交互,而 Tracr 显示所需测量族依赖基底选择。在 GPT-2-small IOI 任务中,名字移动器-负名字移动器交互是三组跨组对中最大的持留预测项。在 Qwen-2.5-7B 上,有限校准使加性拒绝响应图已足够,故持留误差不支持成对提升。
原文摘要 · Abstract (English)
Mechanistic interpretability seeks quantities that models do not expose directly: represented states, component effects, interactions, and responses to interventions. Patching, gradients, Hessian-vector products, and subset interventions provide different measurements under different access assumptions and may target different quantities. We formulate their shared measurement structure as mechanistic tomography: designed measurement for recovering internal mechanisms and intervention effects. For a chosen basis and intervention family, measurements take the form y = Ax + w, where A describes the interventions, x is the target map, and w contains nonlinear response, sampling error, and basis misspecification. This language gives a practical procedure: start with the least costly measurements, test on held-out interventions at the intended scale, calibrate simple mismatch, and expand the measurement family when structured residuals remain. Control provides a demanding validation setting because an estimate that guides an intervention acts as an observer. In a two-HMM model, control error rises with observer error, while target improvement can hide nuisance-state movement. Under forward-only access, sparse aggregate measurements recover a finite-effect map with fewer interventions than coordinate patching. With gradient access, finite probes improve a local attribution map. Lifted measurements and Hessian-vector products recover interactions missed by first-order maps, while Tracr shows that the required family depends on the basis. On GPT-2-small IOI, the Name Mover-Negative Name Mover interaction is the largest held-out predictive term among three tested cross-group pairs. On Qwen-2.5-7B, finite calibration makes an additive refusal-response map adequate, so held-out error does not support pairwise lifting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。