arXiv:2603.04378cs.LGcs.AI2026-03被引 1

通过定向控制对抗方向敏感性,提升智能体系统的鲁棒性与训练稳定性。

Robustness of Agentic AI Systems via Adversarially-Aligned Jacobian Regularization

  • 仅在对抗上升方向约束雅可比范数,避免全局过度平滑
  • 理论证明可容纳更大策略空间,性能下降更小
  • 提供稳定优化步长条件,适合高复杂度多智能体系统

随着大语言模型向自主多智能体系统演进,鲁棒的极小极大训练变得至关重要,但在高度非线性策略导致内层最大化局部曲率极端时仍易失稳。现有通过全局雅可比范数约束的方法过于保守,抑制所有方向的敏感性,带来显著的鲁棒性代价。本文提出对抗对齐雅可比正则化(AAJR),一种沿轨迹对齐的方法,仅在对抗上升方向严格控制敏感性。我们证明,在温和条件下,AAJR 所允许的策略类比全局约束更广,意味着近似误差更小且名义性能损失更低。此外,我们推导出保证优化轨迹上有效平滑性的步长条件,确保内层迭代稳定。这些结果为智能体鲁棒性提供了结构性理论,将极小极大稳定性与全局表达力限制解耦。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) transition into autonomous multi-agent ecosystems, robust minimax training becomes essential yet remains prone to instability when highly non-linear policies induce extreme local curvature in the inner maximization. Standard remedies that enforce global Jacobian bounds are overly conservative, suppressing sensitivity in all directions and inducing a large Price of Robustness. We introduce Adversarially-Aligned Jacobian Regularization (AAJR), a trajectory-aligned approach that controls sensitivity strictly along adversarial ascent directions. We prove that AAJR yields a strictly larger admissible policy class than global constraints under mild conditions, implying a weakly smaller approximation gap and reduced nominal performance degradation. Furthermore, we derive step-size conditions under which AAJR controls effective smoothness along optimization trajectories and ensures inner-loop stability. These results provide a structural theory for agentic robustness that decouples minimax stability from global expressivity restrictions.

智能体系统鲁棒性极小极大训练雅可比正则

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。