arXiv:2608.23725cs.LGcs.NE2026-08

解决深度平衡模型训练中的梯度不稳问题,提升优化可靠性。

Response Renormalization for Critical Deep Equilibrium Models

论文配图:Response Renormalization for Critical Deep Equilibrium Models
图 1 · 摘自论文原文
  • 通过响应归一化技术,局部修正接近奇异的雅可比矩阵引起的梯度放大。
  • 在23类多物理系统中,95%以上场景测试误差仅比精确方法高5%以内。
  • 适合需要稳定训练的复杂隐式模型,如偏微分方程与粒子系统建模。

深度平衡模型(DEQ)通过不动点计算预测,其训练依赖隐式微分,需求解由残差雅可比构建的伴随系统。当该雅可比在损失敏感方向接近奇异时,微小扰动会大幅放大伴随响应,导致梯度过大且敏感,使优化不可靠。本文提出响应归一化(Response Renormalization),在不改变非临界通道的前提下,对部分近极点分母进行提升。集体模式响应归一化(CMR)在低维临界子空间中应用此修正;Phi自适应CMR基于正感应规则计算有界响应质量。我们推导了密集与矩阵自由的集体形式,区分了修改后冻结锚点残差的精确梯度与反向响应代理,并将方法扩展至结构隐式层与向量吸引子(SILVA)。在涵盖偏微分方程、三维场、算子映射、复杂几何与粒子系统的23类多物理家族中,CMR与Phi-CMR在超过98%的静态和95%的瞬态族-种子比较中,测试误差不超过精确隐式微分方法的5%。求解器索引实验显示收敛至静态伴随,物理时间滚动预测仍保持高保真。结果表明,选择性响应归一化可在不全局抑制良好条件敏感性的前提下,有效控制近临界伴随放大,使参数更新更可靠,同时保留学习所需梯度信息。

原文摘要 · Abstract (English)

Deep Equilibrium Models (DEQs) compute predictions from a hidden representation unchanged by the model update. Training through this equilibrium uses implicit differentiation and requires solving an adjoint system built from the residual Jacobian. If this Jacobian is nearly singular along loss-sensitive directions, small perturbations can be strongly amplified in the adjoint response, producing large, highly sensitive gradients that can make optimization unreliable. We introduce Response Renormalization, a backward-pass framework that lifts selected near-pole denominators while leaving unlifted response channels unchanged. Collective Mode Response Renormalization (CMR) applies this correction in a low-dimensional critical subspace, while Phi-adaptive CMR computes a bounded response mass from a positive susceptibility rule. We derive dense and matrix-free collective formulations, distinguish exact gradients of a modified frozen-anchor residual from backward-response surrogates, and extend the construction to Structured Implicit Layers and Vector Attractors (SILVA). Across 23 multiphysics families spanning partial differential equations, three-dimensional fields, operator maps, complex geometries, and particle systems, CMR and Phi-CMR yield test errors no more than five percent higher than those from models trained with exact implicit differentiation in more than 98% of static and 95% of transient family-seed comparisons. Solver-index experiments show convergence toward the static adjoint, while physical-time rollouts retain predictive fidelity under the evaluated conditions. These results demonstrate that selective response renormalization can control near-critical adjoint amplification without globally damping well-conditioned sensitivity. Therefore, the method can make parameter updates more reliable while preserving the useful gradient information needed for learning.

深度平衡隐式微分梯度稳定多物理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。