arXiv:2504.07722cs.LGstat.ME2025-04

提出相对不可忽略性框架,让强化学习在信息不全时仍能稳定收敛。

A Relative Ignorability Framework for Decision-Relevant Observability in Control Theory and Reinforcement Learning

  • 引入相对不可忽略性概念,放宽对观测完整性的要求。
  • 证明在非马尔可夫环境下,标准RL算法仍可收敛。
  • 适合研究鲁棒决策与真实世界智能系统的学者。

序列决策系统常面临数据缺失或不完整问题。经典强化学习理论依赖马尔可夫可观测性假设,但在部分可观测场景下可能失效。因果推断中的不可忽略性概念提供了新视角。本文提出相对不可忽略性这一图因果准则,对基于不完整数据的准确决策进行更精细约束。理论分析与仿真表明,当缺失机制相对于因果估计量具有相对不可忽略性时,即使过程非马尔可夫,标准Q-learning算法仍能收敛。该结果拓展了安全、高效人工智能的理论基础,使其适用于无法获取完整信息的真实环境。

原文摘要 · Abstract (English)

Sequential decision-making systems routinely operate with missing or incomplete data. Classical reinforcement learning theory, which is commonly used to solve sequential decision problems, assumes Markovian observability, which may not hold under partial observability. Causal inference paradigms formalise ignorability of missingness. We show these views can be unified and generalized in order to guarantee Q-learning convergence even when the Markov property fails. To do so, we introduce the concept of relative ignorability. Relative ignorability is a graphical-causal criterion which refines the requirements for accurate decision-making based on incomplete data. Theoretical results and simulations both reveal that non-Markovian stochastic processes whose missingness is relatively ignorable with respect to causal estimands can still be optimized using standard Reinforcement Learning algorithms. These results expand the theoretical foundations of safe, data-efficient AI to real-world environments where complete information is unattainable.

强化学习因果推断决策优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。