arXiv:2605.03288cs.RO2026-05中稿 · ICML

提出神经控制框架,解决形变物体操控中隐式状态依赖历史的梯度传播难题。

Neural Control: Adjoint Learning Through Equilibrium Constraints

论文配图:Neural Control: Adjoint Learning Through Equilibrium Constraints
图 1 · 摘自论文原文
  • 用伴随法在分支依赖的平衡求解序列上反向传播梯度
  • 避免求解器展开,保持前向计算在收敛平衡态上进行
  • 适用于真实形变物体操控,适合长期控制与隐式模型学习

许多物理人工智能任务涉及顺序隐式计算:每一步施加边界控制,通过求解平衡问题获得最终配置。这种设定自然出现在可变形物体操控中,即使是弯曲一根可变形线性物体(DLO)至目标形状,也可能呈现非线性和多稳态特性:相同的边界条件可能因执行历史不同而产生不同构型。与显式转移模型不同,控制到配置的关系是隐式的且依赖历史,导致长时程学习和控制脆弱;而对迭代求解过程进行反向传播则内存与计算开销巨大。我们提出神经控制(Neural Control),一种边界控制框架,通过分支依赖的平衡求解序列传播梯度,而非单一固定点。该方法通过伴随公式微分平衡条件,计算轨迹相关的代理梯度,在不展开求解器的同时,保持前向滚动在收敛平衡态上进行。结合滚动时域延续策略,神经控制将优化重新锚定于实际实现的平衡态,缓解了平衡域切换问题。我们在模拟与真实的DLO操控任务上验证了神经控制的有效性,与SPSA和iCEM对比,展示了其在学习的DEQ风格隐式平衡模型中的适用性。

原文摘要 · Abstract (English)

Many physical AI tasks require sequential implicit computation: at each step, boundary controls are applied, and the resulting configuration is obtained by solving an equilibrium problem. This setting arises naturally in deformable object manipulation, where even bending a deformable linear object (DLO) to a target shape can be nonlinear and multistable: identical boundary conditions may produce different configurations depending on actuation history. Unlike explicit transition models, the control-to-configuration relation is implicit and history-dependent, making long-horizon learning and control brittle; backpropagating through iterative solves is also memory- and compute-intensive. We propose Neural Control, a boundary-control framework that propagates gradients through branch-dependent sequences of equilibrium solves rather than a single fixed point. Neural Control computes trajectory-dependent proxy gradients by differentiating equilibrium conditions with an adjoint formulation, avoiding solver unrolling while keeping forward rollouts on converged equilibria. Combined with receding-horizon continuation, Neural Control re-anchors optimization to realized equilibria and mitigates basin switching. We validate Neural Control on simulated and real DLO manipulation, compare against SPSA and iCEM, and demonstrate applicability to a learned DEQ-style implicit equilibrium model.

物理建模隐式模型梯度传播控制优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。