通过无目标微干预预测大模型训练响应,提升跨家族迁移性能。
Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families
- 从同一检查点分支出四种无目标微干预,捕捉训练响应动态。
- 相比仅用当前能力,误差降低39.4%以上,最优读出方式准确率达75.4%。
- 适用于模型迁移、训练策略优化,尤其适合跨家族分析场景。
基准评分描述了检查点当前的能力,但无法预测其对下一阶段训练的响应。我们通过从同一检查点分支出四个标准化、无目标的微干预,并在统一能力空间中记录其影响,来度量这一缺失状态。结合当前能力,这些响应构成L-State;其脉冲块支持灵活的直接读出与结构保持的算子读出。在局部平滑动态假设下,算子构造可实现端到端跨家族边界,明确体现源与目标家族的坐标异质性。在三家族留一法验证中,两种脉冲读出均使源标准化均方误差相对降低39.4%,并分离出最佳响应与方向估计。在封闭的GLM-4-9B上,直接读出和算子读出分别将误差降低71.8%与78.3%,算子读出将符号平衡准确率从0.366提升至0.754。在封闭Granite-3.1-8B上,直接读出达到RMSE 0.544,经训练适配的动作选择器达0.554,远优于能力单独时的1.172。五家族审计发现算子坐标随任务与家族变化,建模此类偏差可提升回溯轨迹预测性能。因此,无目标干预揭示了当前能力所忽略的训练响应信息,直接与结构化读出覆盖互补的迁移范式。
原文摘要 · Abstract (English)
Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. Together with current capability, these responses form L-State; its pulse block supports a flexible direct readout and a structure-preserving operator readout. Under smooth local dynamics, the operator construction admits an end-to-end cross-family bound with explicit source- and target-family coordinate heterogeneity. In three-family leave-one-family-out development, both pulse readouts reduce source-standardized MSE by 39.4% relative to capability alone, while separating the best response and direction estimates. On sealed GLM-4-9B, the direct and operator readouts reduce MSE by 71.8% and 78.3%, respectively, and the operator readout raises sign balanced accuracy from 0.366 to 0.754. On sealed Granite-3.1-8B, the direct readout reaches RMSE 0.544 and a development-fitted action-wise selector reaches 0.554, compared with 1.172 for capability alone. A five-family audit finds that the operator coordinate varies by action and family, and that modeling these deviations improves retrospective held-trajectory prediction. Target-independent interventions therefore expose training-response information that current capability misses, with direct and structured readouts covering complementary transfer regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。