通过双向神经活动迁移,揭示模型间因果相似性。
Model Alignment Search
- 提出双向神经活动迁移方法,衡量模型功能相似性。
- 实证显示不同任务下循环模型对数字编码方式差异。
- 适用于检测微调模型的对齐偏差,尤其适合生物神经网络对比。
何时可判断两个神经系统以相同方式完成任务?若未因果探测其表征,会遗漏哪些细节?又如何建立双向因果关系?本文提出一种双向转移人工神经网络间神经活动的方法,并用其行为表现作为功能相似性的度量。首先展示该方法可将一个冻结神经网络的行为迁移至另一网络,类似模型拼接;并表明其与表示相似性分析等相关性度量存在本质区别。接着,从实证和理论上证明该方法在特定条件下等价于模型拼接,或聚焦于共享因果信息,且将n个模型间的比较所需矩阵数从二次降低为线性。案例研究显示,该方法能揭示数字相关任务中递归模型的因果编码差异;另验证了微调后的DeepSeek-r1-Qwen-1.5B模型存在对齐问题。最后,引入反事实潜在(CL)辅助目标,增强当一方网络因果不可访问时的因果相关性。结果倡导在神经相似性分析中采用因果方法,并展望未来模型对齐研究路径。
原文摘要 · Abstract (English)
When can we say that two neural systems perform a task in the same way? What nuances do we miss when we fail to causally probe the representations of the systems, and how do we establish bidirectional causal relationships? In this work, we introduce a method that bidirectionally transfers neural activity between artificial neural networks and uses their resulting behavior as a measure of functional similarity. We first show that the method can be used to transfer the behavior from one frozen Neural Network (NN) to another in a manner similar to model stitching, and we show how the method can differ from correlative similarity measures like Representational Similarity Analysis. Next, we empirically and theoretically show how the method can be equivalent to model stitching when desired, or it can take a form that has a more restrictive focus to shared causal information; in both forms, it reduces the number of required matrices for a comparison of n models to be linear in n. We then present a case study on number-related tasks showing that the method can be used to examine specific subtypes of causal information demonstrating that numbers can be encoded differently in recurrent models depending on the task, and we present another case study showing that MAS can reveal misalignment in fine-tuned DeepSeek-r1-Qwen-1.5B models. Lastly, we augment the loss function with a counterfactual latent (CL) auxiliary objective to improve causal relevance when one of the two networks is causally inaccessible (as is often the case in comparisons with biological networks). We use our results to encourage the use of causal methods in neural similarity analyses and to suggest future explorations of network similarity methodology for model misalignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。