发现微调模型合并后会重激活推理痕迹,提出新评估方法与干预机制。
Adapter Merging Reactivates Latent Reasoning Traces: A Mechanism Analysis
- 用无标记答案评估法检测推理痕迹泄漏,避免表面提示干扰。
- 沿特定方向干预模型输出,可显著提升选择题准确率。
- 揭示领域与指令适配器更新方向错位,提出几何感知合并方案。
通过两阶段微调(领域适配后接指令对齐)的大语言模型在适配器合并后可能出现非平凡干扰,包括严格解码下显式推理痕迹的重新出现。本研究在医疗LLM场景中,采用轻量级、可复现的痕迹泄漏与指令遵循行为测量方法分析该现象。除基于标记的代理指标外,我们引入无需标记、仅评估答案的评估方式,并定义不依赖表面标记的正确性方向;沿该方向进行秩1的逻辑空间干预,在足够强的干预强度下可调节决策分布,且优于随机方向控制,提升多项选择准确率。进一步提供层间几何证据,表明领域与指令适配器引发部分错位的更新方向,并展示一个概念验证性的几何感知合并方法,可在模拟场景中降低泄漏并/或提高准确率。结果刻画了痕迹泄漏的边界条件,为更安全的适配器合并提供了实用诊断与干预手段。
原文摘要 · Abstract (English)
Large language models fine-tuned via a two-stage pipeline (domain adaptation followed by instruction alignment) can exhibit non-trivial interference after adapter merging, including the re-emergence of explicit reasoning traces under strict decoding. We study this phenomenon in medical LLM settings using lightweight, reproducible measurements of trace leakage and instruction-following behavior. Beyond marker-based proxies, we introduce a marker-forbidden, answer-only evaluation and define a correctness-based direction that does not rely on surface markers; a rank-1 logit-space intervention along this direction modulates decision distributions and improves multiple-choice accuracy beyond random-direction controls at sufficiently large intervention strength. We further provide layer-wise geometric evidence that domain and instruction adapters induce partially misaligned update directions, and present a proof-of-concept geometry-aware merge that can reduce leakage and/or improve accuracy in a toy setting. Our results characterize boundary conditions of trace leakage and provide practical diagnostics and interventions for safer adapter merging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。