arXiv:2509.15652cs.LG2025-09被引 1

用非凸正则化提升强化学习特征选择精度,有效抑制噪声干扰。

Nonconvex Regularization for Feature Selection in Reinforcement Learning

  • 基于最小二乘时序差分框架,引入非凸稀疏惩罚优化策略评估
  • 在含大量噪声特征场景下,性能显著优于现有方法
  • 理论保证收敛性,适合高维复杂环境中的特征筛选任务

本文提出一种高效的批量特征选择算法,用于强化学习,并具备理论收敛性保障。为降低传统正则化方案固有的估计偏差,首次在经典的最小二乘时序差分(LSTD)框架内,将贝尔曼残差目标与稀疏诱导的非凸投影极小极大凹(PMC)惩罚相结合。由于PMC惩罚具有弱凸性,该问题可视为一类广义非单调包含问题的特例。第二项贡献建立了前向-反射-后向分裂(FRBS)算法对此类问题的新型收敛条件。在基准数据集上的数值实验表明,所提方法在存在大量噪声特征的场景下,显著优于当前最优特征选择方法。

原文摘要 · Abstract (English)

This work proposes an efficient batch algorithm for feature selection in reinforcement learning (RL) with theoretical convergence guarantees. To mitigate the estimation bias inherent in conventional regularization schemes, the first contribution extends policy evaluation within the classical least-squares temporal-difference (LSTD) framework by formulating a Bellman-residual objective regularized with the sparsity-inducing, nonconvex projected minimax concave (PMC) penalty. Owing to the weak convexity of the PMC penalty, this formulation can be interpreted as a special instance of a general nonmonotone-inclusion problem. The second contribution establishes novel convergence conditions for the forward-reflected-backward splitting (FRBS) algorithm to solve this class of problems. Numerical experiments on benchmark datasets demonstrate that the proposed approach substantially outperforms state-of-the-art feature-selection methods, particularly in scenarios with many noisy features.

强化学习特征选择非凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。