arXiv:2609.08618cs.LG2026-09

通过无目标微干预预测大模型训练响应,提升跨家族迁移性能。

Target-Independent Micro-Interventions for Predicting Training Response Across Language-Model Families

  • 从同一检查点分支出四种无目标微干预,捕捉训练响应动态。
  • 相比仅用当前能力,误差降低39.4%以上,最优读出方式准确率达75.4%。
  • 适用于模型迁移、训练策略优化,尤其适合跨家族分析场景。

基准评分描述了检查点当前的能力,但无法预测其对下一阶段训练的响应。我们通过从同一检查点分支出四个标准化、无目标的微干预,并在统一能力空间中记录其影响,来度量这一缺失状态。结合当前能力,这些响应构成L-State;其脉冲块支持灵活的直接读出与结构保持的算子读出。在局部平滑动态假设下,算子构造可实现端到端跨家族边界,明确体现源与目标家族的坐标异质性。在三家族留一法验证中,两种脉冲读出均使源标准化均方误差相对降低39.4%,并分离出最佳响应与方向估计。在封闭的GLM-4-9B上,直接读出和算子读出分别将误差降低71.8%与78.3%,算子读出将符号平衡准确率从0.366提升至0.754。在封闭Granite-3.1-8B上,直接读出达到RMSE 0.544,经训练适配的动作选择器达0.554,远优于能力单独时的1.172。五家族审计发现算子坐标随任务与家族变化,建模此类偏差可提升回溯轨迹预测性能。因此,无目标干预揭示了当前能力所忽略的训练响应信息,直接与结构化读出覆盖互补的迁移范式。

原文摘要 · Abstract (English)

Benchmark scores describe what a checkpoint can do now, but they do not determine how it will respond to the next training episode. We measure this missing state by branching four short, standardized, target-independent micro-interventions from the same checkpoint and recording their effects in a common capability space. Together with current capability, these responses form L-State; its pulse block supports a flexible direct readout and a structure-preserving operator readout. Under smooth local dynamics, the operator construction admits an end-to-end cross-family bound with explicit source- and target-family coordinate heterogeneity. In three-family leave-one-family-out development, both pulse readouts reduce source-standardized MSE by 39.4% relative to capability alone, while separating the best response and direction estimates. On sealed GLM-4-9B, the direct and operator readouts reduce MSE by 71.8% and 78.3%, respectively, and the operator readout raises sign balanced accuracy from 0.366 to 0.754. On sealed Granite-3.1-8B, the direct readout reaches RMSE 0.544 and a development-fitted action-wise selector reaches 0.554, compared with 1.172 for capability alone. A five-family audit finds that the operator coordinate varies by action and family, and that modeling these deviations improves retrospective held-trajectory prediction. Target-independent interventions therefore expose training-response information that current capability misses, with direct and structured readouts covering complementary transfer regimes.

模型预测训练响应跨家族迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。