通过权重变化审计发现医学专用模型提升源于整体更新,非单一模块贡献。
Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

- 采用配对权重增量路径审计,分析通用到医学模型的更新过程。
- 两个模型更新后医疗基准得分提升0.974和1.183,主要由MLP层驱动。
- 结果表明提升非单一组件导致,适合关注模型可解释性的研究者参考。
专用语言模型通常通过终点性能提升来理解:通用模型表现较低,专用模型表现更高,差异被视为专业化的证据。然而,这种分析忽略了模型更新本身的细节。本文提出一种配对权重增量路径审计方法,并应用于两组公开、对齐的通用至医学专用模型检查点:Gemma-3-4B-IT 到 MedGemma-4B-IT,以及 Qwen2.5-7B-Instruct 到 HuatuoGPT-o1-7B。在两组中,完整的解码器侧更新均能有效重建测量到的医疗基准性能提升(0.974 和 1.183 的终点归一化保留率),使每个解码器增量成为审计的合适基础。但性能提升并未被清晰定位。在两组中,MLP 层是表现最强的通用组件家族,但混合跨域变化、10次种子匹配对照与终点锚定回滚的存在,使得无法用单一粗粒度组件解释。因此,审计将更新级重构与组件级解释分离。其结论仅针对纯文本多选题基准性能变化,不涉及临床验证、修复或电路级机制。
原文摘要 · Abstract (English)
Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and the difference is treated as evidence of specialization. This leaves the released update itself largely unexamined. We propose a paired weight-delta path audit and apply it to two public, aligned generalist-to-medical-specialist checkpoint pairs: Gemma-3-4B-IT to MedGemma-4B-IT and Qwen2.5-7B-Instruct to HuatuoGPT-o1-7B. In both pairs, the full decoder-side update strongly reconstructs measured medical benchmark movement (0.974 and 1.183 endpoint-normalized retention), making each decoder delta an appropriate substrate for the audit. Yet the movement is not cleanly localized. MLP is the strongest broad component family in both pairs, but mixed off-domain movements, 10-seed matched controls, and endpoint-anchored rollbacks prevent a unique coarse-family explanation. The audit therefore separates update-level reconstruction from component-level explanation. Its claims concern text-only multiple-choice benchmark movement, not clinical validation, repair, or circuit-level mechanism.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。