arXiv:2602.11729cs.AIcs.LG2026-02被引 9

用跨架构对比发现大模型隐藏行为差异,无需标注数据。

Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs

  • 提出专用特征跨编码器,分离不同模型的独特特征。
  • 无监督发现多个模型的意识形态与版权拒绝机制。
  • 适合安全评估与模型对齐研究者使用。

模型差分通过比较模型内部表示以识别其差异,是发现新模型中关键安全行为的有前景方法。然而,现有方法主要局限于基础模型与微调模型的对比。由于新发布的大型语言模型常采用新颖架构,跨架构差分方法对广泛适用性至关重要。交叉编码器(Crosscoders)是一种可实现跨架构模型差分的技术,但此前仅用于基础模型与微调模型的对比。本文首次将交叉编码器应用于跨架构模型差分,并引入专用特征交叉编码器(DFCs),通过架构改进更好隔离单一模型独有的特征。基于该技术,我们以无监督方式发现了 Qwen3-8B 和 Deepseek-R1-0528-Qwen3-8B 中的中国共产党对齐特征,Llama3.1-8B-Instruct 中的美国例外主义特征,以及 GPT-OSS-20B 中的版权拒绝机制。结果推动了跨架构交叉编码器模型差分作为识别模型间有意义行为差异的有效方法的发展。

原文摘要 · Abstract (English)

Model diffing, the process of comparing models' internal representations to identify their differences, is a promising approach for uncovering safety-critical behaviors in new models. However, its application has so far been primarily focused on comparing a base model with its finetune. Since new LLM releases are often novel architectures, cross-architecture methods are essential to make model diffing widely applicable. Crosscoders are one solution capable of cross-architecture model diffing but have only ever been applied to base vs finetune comparisons. We provide the first application of crosscoders to cross-architecture model diffing and introduce Dedicated Feature Crosscoders (DFCs), an architectural modification designed to better isolate features unique to one model. Using this technique, we find in an unsupervised fashion features including Chinese Communist Party alignment in Qwen3-8B and Deepseek-R1-0528-Qwen3-8B, American exceptionalism in Llama3.1-8B-Instruct, and a copyright refusal mechanism in GPT-OSS-20B. Together, our results work towards establishing cross-architecture crosscoder model diffing as an effective method for identifying meaningful behavioral differences between AI models.

模型差分大模型安全无监督学习跨架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。