不训练就能提升指令模型的探索能力,让回答更灵活。
Timber: Training-free Instruct Model Refining with Base via Effective Rank
- 通过微调权重差异,将指令模型部分还原为基础模型以增强探索性。
- 在Llama和Qwen系列上,Pass@k指标显著提升,尤其在复杂任务上表现更好。
- 无需训练,适合希望优化现有模型而不增加成本的研究者。
后训练阶段通常被认为只是表面调整,本文从权重层面提供新量化证据:有效秩(eRank)几乎不变。然而这种表面性也带来关键权衡——提升利用能力的同时限制了探索能力。为此,我们提出Timber,一种简单有效的免训练方法,通过精细调整权重差异,将指令模型部分回退至对应基础模型,从而增强其探索能力并保持原有利用性能。在Llama与Qwen系列上的大量实验表明,Timber持续提升原始指令模型的表现,尤其在Pass@k指标上效果显著。研究揭示了后训练阶段在权重层面的新见解,并提供了无需训练即可优化指令模型的实际策略。
原文摘要 · Abstract (English)
Post-training, which elicits a pretrained Base model into the corresponding Instruct model, is widely considered to be superficial. In this work, we first reinforce this hypothesis by providing novel quantitative evidence from the weight level that the effective rank (eRank) remains negligibly changed. However, this superficiality also suffers a critical trade-off, improving the exploitation capabilities at the cost of limiting its exploration. To tackle this issue, we propose Timber, a simple yet effective training-free method that enhances the exploration capability of the Instruct model while preserving its exploitation. The key insight is to partially revert Instruct towards the paired Base model by subtle yet targeted refinement of the weight deltas. Extensive experiments on Llama and Qwen series demonstrate that Timber consistently improves vanilla Instruct models, particularly on Pass@k performance. Our findings offer new insights into the post-training stage at the weight level and practical strategies to refine the Instruct model without training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。