验证模型是否在符号域间传递操作级因果结构。
A Pre-Specified Construction-Confirmation Test of Operation-Level Causal Transfer Across Finite Isomorphic Symbolic Domains
- 设计输入特异性干预测试,分离构建与验证阶段。
- 单条提示路径在双实验中均通过显著性检验(p=0.000198)。
- 结果仅限特定路径,适用于模型内部机制分析。
行为准确率、线性可解性和成功激活干预本身并不能证明模型将操作级结构从一个符号域传递到另一个域。本文在有限同构状态空间中提出更具体的问题:若分别估计两个操作在源输入中的隐藏状态差异,将其加至映射后的目标输入,能否使模型趋近对应的目标答案?实验对比了该输入特定干预与错误操作、同范数随机及无操作控制组,并将候选构建与独立确认分割。在冻结的Qwen2.5-7B-Instruct模型第20–21层,一条预设且冻结的候选路径——透明 | integer_mod16--letters16 | successor→predecessor——通过了两个PyVene分割;其确认交集-并集p值为0.000198,36族霍尔姆校正后p值为0.006943。后续使用NNsight 0.7.0的实验,同样在确认前预设并冻结,仅测试该选定提示路径,未重新选择候选或层数,成功复现全部12个确认效应估计值、置信区间及精确符号翻转p值,其36族霍尔姆校正后p值为0.007141。因此结果局限于单一提示路径与候选,跨两种干预实现、一个模型版本和一层区间得到复现。结论不支持跨模型泛化、全族后端独立性、通用域迁移或代数不变性。
原文摘要 · Abstract (English)
Behavioral accuracy, linear decodability, and successful activation interventions do not by themselves show that a model carries an operation-level structure from one symbolic domain to another. We ask a narrower question in finite isomorphic state spaces: if the hidden-state difference between two operations is estimated separately for each source input, does adding that difference to a mapped recipient input move the model toward the corresponding recipient answer? The design compares this input-specific intervention with wrong-operation, norm-matched random, and no-op controls, and separates candidate construction from an independently isolated confirmation split. On a frozen Qwen2.5-7B-Instruct model at layers 20--21, one route--domain--operation candidate from a family pre-specified and frozen before confirmation access, transparent | integer_mod16--letters16 | successor->predecessor, passed both PyVene splits; its confirmation intersection--union p-value was 0.000198 and its 36-family Holm-adjusted p-value was 0.006943. A subsequent NNsight 0.7.0 experiment, pre-specified and frozen before its confirmation access, tested only this selected prompt route, without candidate or layer reselection. It reproduced all 12 confirmation effect estimates, confidence intervals, and exact sign-flip p-values numerically; its 36-family Holm-adjusted p-value was 0.007141. The result is therefore limited to one prompt route and one candidate, replicated across two intervention implementations on one model revision and one layer interval. It does not establish cross-model generalization, full-family backend independence, domain-general transfer, or algebraic invariance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。