从异质数据中学习个性化最优策略,提升复杂环境下的决策效果
Reinforcement Learning for Individual Optimal Policy from Heterogeneous Data
- 引入个体隐变量建模异质性,实现个性化策略学习
- 在弱部分覆盖假设下,平均后悔率收敛速度快于现有方法
- 适用于有差异个体的现实场景,如医疗、推荐系统
离线强化学习旨在利用预先收集的数据,在动态环境中寻找最大化期望总回报的最优策略。从异质数据中学习是离线强化学习的核心挑战之一。传统方法通常基于单个或同质批次数据学习对所有个体通用的最优策略,可能导致异质群体中的次优表现。本文提出一种针对异质时平稳马尔可夫决策过程的个体化离线策略优化框架。通过引入包含个体隐变量的异质模型,可高效估计个体级Q函数;提出的惩罚性悲观个性化策略学习(P4L)算法,在行为策略满足弱部分覆盖假设下,保证了平均后悔率的快速收敛。仿真研究与真实数据应用均表明,该方法性能显著优于现有方法。
原文摘要 · Abstract (English)
Offline reinforcement learning (RL) aims to find optimal policies in dynamic environments in order to maximize the expected total rewards by leveraging pre-collected data. Learning from heterogeneous data is one of the fundamental challenges in offline RL. Traditional methods focus on learning an optimal policy for all individuals with pre-collected data from a single episode or homogeneous batch episodes, and thus, may result in a suboptimal policy for a heterogeneous population. In this paper, we propose an individualized offline policy optimization framework for heterogeneous time-stationary Markov decision processes (MDPs). The proposed heterogeneous model with individual latent variables enables us to efficiently estimate the individual Q-functions, and our Penalized Pessimistic Personalized Policy Learning (P4L) algorithm guarantees a fast rate on the average regret under a weak partial coverage assumption on behavior policies. In addition, our simulation studies and a real data application demonstrate the superior numerical performance of the proposed method compared with existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。