提出安全反学习框架,可移除恶意数据影响而不需重训练
Safe-RULE: Safe Reinforcement UnLEarning

- 通过显式考虑任务性能与安全约束实现反学习
- 在基准任务中有效提升对数据投毒攻击的安全性
- 适合用于机器人等高安全性场景的防御
离线安全强化学习(Safe RL)可在无需在线交互的情况下学习策略,适用于机器人等安全关键系统。然而,其依赖静态数据集的特性使其易受数据投毒攻击——攻击者注入恶意样本会破坏安全性和诱导不安全策略行为。本文提出一种新型学习范式:安全反学习(Safe-RULE),作为防御框架,在不重新训练或不访问原始训练环境的前提下,消除恶意数据的影响。我们进一步将强化反学习扩展至离线安全强化学习,明确在反学习过程中同时考虑任务性能与安全约束。在多个基准安全强化学习任务上的实验表明,该方法能有效提升对数据投毒攻击的安全性。
原文摘要 · Abstract (English)
Offline safe reinforcement learning (Safe RL) enables policy learning without online interactions, making it suitable for safety-critical systems such as robotics systems. However, its reliance on static datasets exposes offline Safe RL to data poisoning attacks, where adversaries inject malicious samples that compromise safety and induce unsafe policy behavior. In this work, we propose a new learning paradigm, named safe reinforcement unlearning (Safe-RULE), used as a defense framework to remove the influence of poisoned data without retraining from scratch or requiring access to the original training environment. We further extend reinforcement unlearning to offline Safe RL by explicitly accounting for both task performance and safety constraints during the unlearning process. Experiments across benchmark Safe RL tasks demonstrate that our approach effectively enhances safety performance against data poisoning attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。