arXiv:2511.04454cs.CEcs.LG2025-11被引 1

用凸松弛法高效拟合强化学习模型,加速行为数据分析。

Fitting Reinforcement Learning Model to Behavioral Data under Bandits

  • 基于凸松弛构建通用优化框架,适配多种强化学习模型。
  • 数值实验显示性能接近顶尖方法,计算速度显著提升。
  • 开源Python工具包,无需优化知识即可直接使用。

本文研究在多臂赌博机环境下,将强化学习(RL)模型拟合到给定行为数据的问题。这类模型近年来广泛用于刻画人类和动物的决策行为。我们提出了一个通用的数学优化问题框架,适用于科研中常见的多种RL模型,并对其凸性性质进行了详细理论分析。基于理论结果,提出一种基于凸松弛与优化的新解法。在多个模拟和真实世界赌博机环境中,该方法与文献中的基准方法进行比较,结果显示其性能可媲美当前最优方法,同时大幅降低计算时间。此外,我们还发布了开源Python包,使研究人员能无需掌握凸优化知识,直接应用于自身数据集分析。

原文摘要 · Abstract (English)

We consider the problem of fitting a reinforcement learning (RL) model to some given behavioral data under a multi-armed bandit environment. These models have received much attention in recent years for characterizing human and animal decision making behavior. We provide a generic mathematical optimization problem formulation for the fitting problem of a wide range of RL models that appear frequently in scientific research applications. We then provide a detailed theoretical analysis of its convexity properties. Based on the theoretical results, we introduce a novel solution method for the fitting problem of RL models based on convex relaxation and optimization. Our method is then evaluated in several simulated and real-world bandit environments to compare with some benchmark methods that appear in the literature. Numerical results indicate that our method achieves comparable performance to the state-of-the-art, while significantly reducing computation time. We also provide an open-source Python package for our proposed method to empower researchers to apply it in the analysis of their datasets directly, without prior knowledge of convex optimization.

强化学习行为建模凸优化数据分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。