为强化学习中的分布偏移提供统一因果分类,揭示其来源与机制。
A Unified Causal-Origin Taxonomy of Distributional Shifts in Reinforcement Learning

- 从环境-智能体交互过程出发,分解出状态、观测、策略等五类结构性因素
- 区分内部(智能体驱动)与外部(环境驱动)偏移,提出显式、隐式、混合三类时间边界偏移
- 统一了分布内/外泛化与非平稳性问题,适合研究鲁棒性与适应性的学者
强化学习系统在运行条件与训练阶段不同时性能下降,源于数据生成过程的分布偏移。这类偏移既可能出现在训练与评估之间(如分布内/外泛化),也可能发生在环境动态随时间演化的非平稳场景中。然而,现有研究多关注缓解方法,忽视了偏移在智能体-环境交互中的因果起源。本文构建一个统一的因果起源分类体系,将经典监督学习中的数据分布偏移原则迁移至强化学习,基于部分可观测马尔可夫决策过程(POMDP),将交互过程分解为状态分布、观测过程、策略、奖励和转移动态,以及偏移的时间边界。该分类区分了内部(智能体驱动)与外部(环境驱动)分布偏移,并通过时间边界视角刻画显式、隐式与混合偏移。此框架将分布内/外泛化与非平稳性统一为底层过程的结构化变化。我们还引入评估框架,通过性能下降与恢复指标量化偏移影响与适应能力。该工作为强化学习在分布偏移下的鲁棒性分析提供了系统性基础。
原文摘要 · Abstract (English)
Reinforcement learning (RL) systems often degrade when operating conditions differ from those previously encountered, reflecting distributional shifts in the underlying data-generating process. Such shifts may occur between training and evaluation, as in In-Distribution (ID) and Out-of-Distribution (OOD) generalization, or within non-stationary settings where environment dynamics evolve over time. However, the formal relationship between these views remains unclear, and existing work mainly focuses on mitigation rather than the causal origin of shift within the agent-environment interaction. This work develops a unified causal-origin taxonomy that characterizes sources of distributional shift in RL and relates ID/OOD generalization to non-stationary settings. We transfer the classical dataset-shift principle from supervised learning to RL by reformulating distributional shift in terms of the generative interaction process. Using a Partially Observable Markov Decision Process (POMDP), we decompose the interaction into structural components, including the state distribution, observation process, policy, reward, and transition dynamics, together with the shifted-time boundary. The proposed taxonomy distinguishes internal (agent-driven) and external (environment-driven) distributional shifts. The shifted-time boundary perspective further characterizes explicit, implicit, and hybrid shifts. This formulation unifies ID/OOD generalization and non-stationarity as structured changes in the underlying process. We also introduce an evaluation framework for measuring shift impact and adaptation through performance degradation and recovery metrics. By grounding distributional shift in the causal-origin structure of RL, this work supports systematic analysis of robustness under distributional shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。