arXiv:2608.01130cs.LG2026-08

研究轨迹更新如何提升决策效果,揭示了何时更新能同时优化训练损失和实际任务表现。

When Do Surrogate Updates Improve Decisions? A Local Theory of Trajectory-Wise Transfer

  • 用可学习性与决策效用定义轨迹更新的两种影响
  • 发现梯度对齐是实现有效转移的关键条件
  • 适用于强化学习与大模型微调场景

大量模型在训练时使用轨迹损失更新,但评估依赖下游任务奖励,存在更新目标与实际决策效果不一致的问题。本文固定检查点与更新空间,将轨迹带来的总体代理损失下降定义为可学习性,决策风险下降定义为决策效用。理论分析表明:一、单步转移误差由梯度方向偏差与曲率共同决定;二、当代理梯度与决策梯度正向共线时,所有可访问方向均实现普适一阶转移;三、校准差距限制基于可学习性的轨迹选择代价,候选差值修正进一步收紧该界;四、在嵌套更新空间中存在近似-校准权衡。网格世界与大模型后训练实验验证了上述预测。

原文摘要 · Abstract (English)

A broad range of models face the mismatch where they are updated through trajectory losses but are evaluated by downstream task reward. Here, a trajectory is a training instance that induces a surrogate loss whose reduction might not track the model's decision utility update. Theoretically, we ask when one step of trajectory training reduces both population surrogate loss and decision risk, and how transfer accumulates along repeated updates. To formalize this, we first fix a checkpoint and a restricted update space, and define the reductions in population surrogate risk and decision risk induced by a trajectory as its learnability and decision utility, respectively. On this basis, our theory yields four main results. First, a one-step transfer bound separates their discrepancy into first-order gradient misalignment after nonnegative calibration and second-order curvature; and a pathwise extension accumulates the same terms over repeated updates. Second, when the accessible surrogate gradient is nonzero, universal first-order transfer over every accessible direction holds exactly when the accessible surrogate and decision gradients are positively collinear. Third, the calibration gap bounds the decision regret of learnability-based trajectory selection, while a candidate-difference refinement tightens this guarantee by retaining only directions that affect pairwise rankings. Finally, we establish an approximation--calibration trade-off across nested update spaces. Controlled gridworld and LLM post-training experiments yield results consistent with our predictions.

强化学习轨迹优化决策效用梯度对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。