arXiv:2509.18792cs.CL2025-09EMNLP被引 1

通过模型差异分析,揭示微调后大模型性能提升的内在机制。

Beyond the Leaderboard: Understanding Performance Disparities in Large Language Models via Model Diffing

  • 用模型差分技术对比两个版本模型的潜在表征差异。
  • 发现增强版在安全、多语言和指令遵循上分别提升32.8%、43.8%、151.7%。
  • 适合关注模型可解释性与微调效果归因的研究者。

随着微调成为提升大语言模型(LLMs)的主要方法,理解其过程中的变化愈发重要。传统基准测试常无法解释为何一个模型优于另一个。本文采用模型差分这一机制可解释性方法,分析Gemma-2-9b-it与其SimPO增强变体之间的具体能力差异。通过crosscoders识别并分类区分两模型的潜在表征。结果显示,SimPO增强了安全机制(+32.8%)、多语言能力(+43.8%)和指令遵循能力(+151.7%),同时降低了对自我引用(-44.1%)和幻觉管理(-68.5%)的关注。该分析表明,模型差分能提供超越排行榜指标的细粒度洞察,将性能差距归因于具体的机制能力,为模型比较提供透明且精准的框架。

原文摘要 · Abstract (English)

As fine-tuning becomes the dominant paradigm for improving large language models (LLMs), understanding what changes during this process is increasingly important. Traditional benchmarking often fails to explain why one model outperforms another. In this work, we use model diffing, a mechanistic interpretability approach, to analyze the specific capability differences between Gemma-2-9b-it and a SimPO-enhanced variant. Using crosscoders, we identify and categorize latent representations that differentiate the two models. We find that SimPO acquired latent concepts predominantly enhance safety mechanisms (+32.8%), multilingual capabilities (+43.8%), and instruction-following (+151.7%), while its additional training also reduces emphasis on model self-reference (-44.1%) and hallucination management (-68.5%). Our analysis shows that model diffing can yield fine-grained insights beyond leaderboard metrics, attributing performance gaps to concrete mechanistic capabilities. This approach offers a transparent and targeted framework for comparing LLMs.

模型差分可解释性微调分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。