arXiv:2601.18952cs.LGstat.ME2026-01被引 1

用希尔伯特空间嵌入方法高效评估多维奖励下的强化学习策略

Vector-Valued Distributional Reinforcement Learning Policy Evaluation: A Hilbert Space Embedding Approach

  • 将概率分布映射到希尔伯特空间,用积分概率度量替代难计算的Wasserstein距离
  • 在多维连续状态动作空间中实现稳定离线策略评估,理论保证收敛性
  • 适合复杂决策场景中的风险评估,尤其适用于多维奖励系统

我们提出一种(离线)多维分布强化学习框架KE-DRL,利用希尔伯特空间映射估计目标策略下多维价值分布的核均值嵌入。在该设置中,状态-动作变量为多维连续型。通过核均值嵌入将概率测度映射至再生核希尔伯特空间,我们的方法用积分概率度量取代了计算困难的Wasserstein距离,从而在高维状态-动作空间和多维奖励设置中实现高效估计。理论上,我们针对马蒂恩核族提出了分布贝尔曼算子的压缩性质,并提供了统一收敛性保证。模拟与实证结果表明,在李普希茨连续性和核有界性的温和假设下,方法具备稳健的离线策略评估能力并能准确恢复核均值嵌入,凸显嵌入方法在复杂现实决策与风险评估中的潜力。

原文摘要 · Abstract (English)

We propose an (offline) multi-dimensional distributional reinforcement learning framework (KE-DRL) that leverages Hilbert space mappings to estimate the kernel mean embedding of the multi-dimensional value distribution under a proposed target policy. In our setting, the state-action variables are multi-dimensional and continuous. By mapping probability measures into a reproducing kernel Hilbert space via kernel mean embeddings, our method replaces Wasserstein metrics with an integral probability metric. This enables efficient estimation in multi-dimensional state-action spaces and reward settings, where direct computation of Wasserstein distances is computationally challenging. Theoretically, we establish contraction properties of the distributional Bellman operator under our proposed metric involving the Matern family of kernels and provide uniform convergence guarantees. Simulations and empirical results demonstrate robust off-policy evaluation and recovery of the kernel mean embedding under mild assumptions, namely, Lipschitz continuity and boundedness of the kernels, highlighting the potential of embedding-based approaches in complex real-world decision-making scenarios and risk evaluation.

分布强化学习核方法策略评估多维奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。