arXiv:2506.08340cs.LG2025-06被引 1

将策略视为自主动力系统,直接优化参数而不依赖强化学习。

Dynamical System Optimization

  • 把策略看作自主动力系统,统一优化参数
  • 算法能计算出与策略梯度相同的量
  • 适合生成模型调优、行为克隆等场景

我们提出一种以参数化策略为中心的优化框架:一旦指定策略,控制权即转移给该策略,形成一个自主动力系统。因此,可直接优化策略参数,无需再参考控制或动作,也无需使用近似动态规划或强化学习的复杂机制。本文推导出在自主系统层面更简单的算法,证明其可计算与策略梯度、海森矩阵、自然梯度、近端方法相同的量。类似近似策略迭代和离线学习的变体也适用。由于策略参数与其他系统参数处理方式一致,同一套算法可用于行为克隆、机制设计、系统辨识、状态估计器学习。生成式AI模型的调优不仅可行,且在概念上比强化学习更贴近本框架。

原文摘要 · Abstract (English)

We develop an optimization framework centered around a core idea: once a (parametric) policy is specified, control authority is transferred to the policy, resulting in an autonomous dynamical system. Thus we should be able to optimize policy parameters without further reference to controls or actions, and without directly using the machinery of approximate Dynamic Programming and Reinforcement Learning. Here we derive simpler algorithms at the autonomous system level, and show that they compute the same quantities as policy gradients and Hessians, natural gradients, proximal methods. Analogs to approximate policy iteration and off-policy learning are also available. Since policy parameters and other system parameters are treated uniformly, the same algorithms apply to behavioral cloning, mechanism design, system identification, learning of state estimators. Tuning of generative AI models is not only possible, but is conceptually closer to the present framework than to Reinforcement Learning.

优化框架动力系统生成模型策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。