用不确定性建模提升离线强化学习在各种数据污染下的鲁棒性
Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data Corruptions

- 将各类数据污染视为动作价值函数的不确定性,通过贝叶斯推断捕捉
- 利用熵度量区分脏数据与干净数据,降低脏数据影响
- 在多种污染场景下显著优于现有方法,适合高噪声或对抗环境
现实世界中的离线数据集常因传感器故障或恶意攻击而存在数据污染(如噪声或对抗攻击)。尽管离线强化学习已取得进展,但现有方法在面对由多样化污染数据(如污染状态、动作、奖励和动态)引发的高不确定性时仍表现不佳,导致在干净环境中的性能下降。为此,我们提出一种新型鲁棒变分贝叶斯离线强化学习方法(TRACER),首次引入贝叶斯推断来通过离线数据捕捉不确定性以增强鲁棒性。TRACER首先将所有污染建模为动作价值函数的不确定性,随后利用全部离线数据作为观测,在贝叶斯框架下近似动作价值函数的后验分布。其关键特性在于:通过熵度量可区分脏数据与干净数据,因为污染数据通常引发更高不确定性与熵值。基于该度量,TRACER可调节脏数据相关的损失,降低其影响,从而提升在干净环境中的性能。实验表明,TRACER在单一及同时存在多种数据污染的场景下均显著优于多个先进方法。
原文摘要 · Abstract (English)
Real-world offline datasets are often subject to data corruptions (such as noise or adversarial attacks) due to sensor failures or malicious attacks. Despite advances in robust offline reinforcement learning (RL), existing methods struggle to learn robust agents under high uncertainty caused by the diverse corrupted data (i.e., corrupted states, actions, rewards, and dynamics), leading to performance degradation in clean environments. To tackle this problem, we propose a novel robust variational Bayesian inference for offline RL (TRACER). It introduces Bayesian inference for the first time to capture the uncertainty via offline data for robustness against all types of data corruptions. Specifically, TRACER first models all corruptions as the uncertainty in the action-value function. Then, to capture such uncertainty, it uses all offline data as the observations to approximate the posterior distribution of the action-value function under a Bayesian inference framework. An appealing feature of TRACER is that it can distinguish corrupted data from clean data using an entropy-based uncertainty measure, since corrupted data often induces higher uncertainty and entropy. Based on the aforementioned measure, TRACER can regulate the loss associated with corrupted data to reduce its influence, thereby enhancing robustness and performance in clean environments. Experiments demonstrate that TRACER significantly outperforms several state-of-the-art approaches across both individual and simultaneous data corruptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。