用编辑操作度量检测强化学习中的环境漂移
Detecting Model Drifts in Non-Stationary Environment Using Edit Operation Measures
- 通过状态动作轨迹的编辑操作度量分布变化
- 在不同噪声下有效区分漂移与非漂移场景
- 适合需要实时监控环境变化的RL应用
强化学习(RL)代理通常假设环境动态是平稳的。然而在医疗、机器人和金融等实际应用中,转移概率或奖励函数可能发生演变,导致模型漂移。本文提出一种新框架,通过分析代理行为序列的分布变化来检测此类漂移。具体而言,我们引入一组基于编辑操作的度量方法,量化在平稳与扰动条件下生成的状态-动作轨迹之间的偏差。实验表明,这些度量即使在不同噪声水平下也能有效区分漂移与非漂移情形,为非平稳强化学习环境中的漂移检测提供了一种实用工具。
原文摘要 · Abstract (English)
Reinforcement learning (RL) agents typically assume stationary environment dynamics. Yet in real-world applications such as healthcare, robotics, and finance, transition probabilities or reward functions may evolve, leading to model drift. This paper proposes a novel framework to detect such drifts by analyzing the distributional changes in sequences of agent behavior. Specifically, we introduce a suite of edit operation-based measures to quantify deviations between state-action trajectories generated under stationary and perturbed conditions. Our experiments demonstrate that these measures can effectively distinguish drifted from non-drifted scenarios, even under varying levels of noise, providing a practical tool for drift detection in non-stationary RL environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。