打破深度Q学习独立性假设,给出依赖数据下的样本保障。
Beyond the Independence Assumption: Finite-Sample Guarantees for Deep Q-Learning under $τ$-Mixing
- 将DQN更新建模为τ-混合数据上的非参数回归问题。
- 发现时间依赖导致统计速率下降,有效样本量减少。
- 在标准环境验证了数据依赖性,支持理论框架。
深度Q学习的有限样本分析通常假设重放缓冲区中的数据是独立的,尽管它们来自具有时间依赖性的状态动作轨迹。本文在显式依赖条件下研究深度Q网络(DQN)算法,将用于更新网络的小批量数据建模为τ-混合过程。我们证明,在特定轨迹依赖条件和采样机制下,该假设成立。基于此,我们将全连接ReLU架构的DQN统计分析拓展至依赖数据。将每次更新视为τ-混合观测的非参数回归问题,并在此依赖结构下推导出有限样本风险界。结果表明,时间依赖性通过在速率指数中引入额外的维度惩罚,导致统计速率退化,反映了τ-混合数据的有效样本量减少。此外,从这些风险界中推导出DQN在τ-混合条件下的样本复杂度。最后,我们在标准Gymnasium环境中进行了实证验证,发现独立性假设被系统性违反,且重放缓冲采样产生近似指数衰减的相关性,支持了理论框架。
原文摘要 · Abstract (English)
Finite-sample analyses of deep Q-learning typically treat replayed data as independent, even though it is sampled from temporally dependent state-action trajectories. We study the Deep Q-networks (DQN) algorithm under explicit dependence by modelling the minibatches used for updating the network as $τ$-mixing. We show that this assumption holds under certain dependence conditions on the underlying trajectories and the mechanism used to sample minibatches. Building on this observation, we extend statistical analyses of DQN with fully connected ReLU architectures to dependent data. We formulate each update as a nonparametric regression problem with $τ$-mixing observations and derive finite-sample risk bounds under this dependence structure. Our results show that temporal dependence leads to a degradation in the statistical rate by inducing an additional dimensionality penalty in the rate exponent, reflecting the reduced effective sample size of $τ$-mixing data. Moreover, we derive the sample complexity of DQN under $tau$-mixing from these risk bounds. Finally, we empirically demonstrate on standard Gymnasium environments that the independence assumption is systematically violated and that replay sampling yields approximately exponentially decaying correlations, supporting our theoretical framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。