arXiv:2605.10658cs.LG2026-05

零阶优化更少遗忘,因随机形变能保留关键记忆结构。

Why Zeroth-Order Adaptation May Forget Less: A Randomized Shaping Theory

  • 用随机形变分析揭示零阶优化的内在保护机制。
  • 在有限查询下,零阶比一阶遗忘减少约20%以上(实验验证)。
  • 适合关注持续学习中记忆稳定性与模型适应性的研究者。

持续学习要求在不损害已有能力的前提下适应新任务。近期研究表明,低查询量的零阶(ZO)优化比一阶(FO)方法遗忘更少,但传统将零阶视为噪声一阶估计的观点无法解释这一现象。本文提出局部随机梯度形变分析:有限差分暴露的原始形状与一阶方向均值对齐,而范数匹配的比较器固定了期望平方更新范数。在此控制条件下,遗忘取决于适应形状对保留曲率的暴露程度。对于范数匹配的零阶,期望形变后的保留曲率满足精确恒等式,保持各向同性保留基线,仅收缩各向异性成分。将此恒等式投影到输入梯度上,得到可观测的 FO-ZO 二次遗忘差距:当一阶方向具有高于平均的保留曲率时,零阶可减少遗忘,降幅为该曲率超额部分的查询依赖比例。实际有限查询计算分离了均值机制与单批次采样及平滑扰动。作为算法迁移,RISE 将校准后的零阶形变应用于参数块内的精确一阶梯度。其目标是平衡稳定性和可塑性:随机形变降低一阶的保留暴露,精确梯度消除有限差分零阶的有限平滑偏差,块内采样则在一次梯度计算后提供多个局部形变方向。块对角曲率分析分离了均值步损伤与中心随机暴露,表明块间耦合、局部形变诊断可明确该精确梯度转移最可能显现的位置。

原文摘要 · Abstract (English)

Continual learning requires new-task adaptation without damaging previously acquired capabilities. Recent forward-pass and zeroth-order (ZO) results show that low-query adaptation may retain better than first-order (FO) descent, but the usual view of ZO as noisy FO estimation does not explain why. We give a local randomized gradient-shaping analysis: finite differences expose a raw shape that is mean-aligned with FO, while the norm-matched comparator fixes the expected squared adaptation norm. Under this controlled comparison, forgetting depends on how the adaptation shape exposes retention curvature. For norm-matched ZO, the expected shaped retention curvature obeys an exact identity that preserves the isotropic retention floor while contracting only the anisotropic component. Projecting this identity onto the incoming gradient yields the observable FO--ZO quadratic forgetting gap: ZO improves mean forgetting precisely when the FO direction has above-average retention curvature, by a query-dependent fraction of that curvature excess. A practical finite-query accounting separates the mean mechanism from one-batch sampling and smoothing perturbations. As an algorithmic transfer, RISE applies the calibrated ZO shape to exact FO gradients inside parameter blocks. Its target is a stability--plasticity tradeoff: randomized shaping may reduce the retention exposure paid by FO, exact gradients remove finite-smoothing bias from finite-difference ZO, and blockwise sampling supplies many local shaping directions after one gradient computation. The blockwise analysis separates mean-step damage from centered random exposure, showing how block-diagonal curvature, cross-block coupling, and local shaping diagnostics specify where this exact-gradient transfer is most likely to be visible.

持续学习零阶优化记忆保留随机形变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。