arXiv:2603.26464cs.LGmath.DS2026-03

用柯尔莫哥洛夫算子自动学特征,让强化学习更省心。

Automatic feature identification in least-squares policy iteration using the Koopman operator framework

  • 用柯尔莫哥洛夫自编码器自动学特征,无需预设
  • 在链式行走和倒立摆任务中收敛效果与传统方法相当
  • 适合缺乏先验特征知识的强化学习场景

本文提出基于柯尔莫哥洛夫算子框架的最小二乘策略迭代算法(KAE-LSPI),通过将最小二乘固定点逼近方法重新表述为扩展动态模式分解(EDMD),实现特征的自动学习。该方法旨在解决线性强化学习中特征或核函数选择缺乏系统性的问题。我们在随机链式行走和倒立摆控制问题上对比了KAE-LSPI与经典最小二乘策略迭代(LSPI)及基于核的最小二乘策略迭代(KLSPI)方法。与以往工作不同,本方法无需预先设定特征或核函数。实验表明,KAE技术学习到的特征数量合理,且收敛至最优或近优策略的效果与另外两种方法相当。

原文摘要 · Abstract (English)

In this paper, we present a Koopman autoencoder-based least-squares policy iteration (KAE-LSPI) algorithm in reinforcement learning (RL). The KAE-LSPI algorithm is based on reformulating the so-called least-squares fixed-point approximation method in terms of extended dynamic mode decomposition (EDMD), thereby enabling automatic feature learning via the Koopman autoencoder (KAE) framework. The approach is motivated by the lack of a systematic choice of features or kernels in linear RL techniques. We compare the KAE-LSPI algorithm with two previous works, the classical least-squares policy iteration (LSPI) and the kernel-based least-squares policy iteration (KLSPI), using stochastic chain walk and inverted pendulum control problems as examples. Unlike previous works, no features or kernels need to be fixed a priori in our approach. Empirical results show the number of features learned by the KAE technique remains reasonable compared to those fixed in the classical LSPI algorithm. The convergence to an optimal or a near-optimal policy is also comparable to the other two methods.

强化学习特征学习柯尔莫哥洛夫

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。