arXiv:2607.21111cs.LGcs.AI2026-07

提出轨迹级遗忘评估基准,解决离线强化学习中数据删除效果难衡量的问题。

TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning

论文配图:TOUR: A Trajectory-Level Unlearning Benchmark for Offline Reinforcement Learning
图 1 · 摘自论文原文
  • 构建包含对照组与多攻击的轨迹级遗忘评估框架。
  • 发现常见删除方法在不同环境中表现不一,且单一评分会高估删除效果。
  • 适合关注数据隐私与模型安全性的强化学习研究者使用。

离线强化学习代理基于固定行为轨迹训练,当需移除特定数据时,轨迹级删除至关重要。但评估删除效果困难,因较低的成员得分可能反映轨迹删除、残留记忆或策略崩溃。本文提出轨迹级遗忘与去记忆评估基准TOUR,整合轨迹划分、匹配非成员对照、重训参考、保留性能锚点及多攻击隐私审计。在D4RL运动控制任务和探索性AntMaze扩展中,TOUR显示常见删除基线具有环境依赖的隐私-效用行为:重训与微调常比均匀GA+Refit提供更强的保留效用参考;尽管TrajDeleter仍为有效参照,但在相同审计下并非始终更优。参考模型、阈值、偏差、等价性、动作误差、表示基及查询受限攻击进一步表明,单一似然成员得分可能夸大删除质量。因此,在所评估设置下,离线强化学习遗忘结论无法通过单一分数稳定判断,其结论依赖于匹配非成员构造、重训相对校准、攻击家族、保留效用及诊断架构或组件层面证据。

原文摘要 · Abstract (English)

Offline Reinforcement Learning (RL) agents are trained on fixed behavioral trajectories, which makes trajectory-level deletion important when selected data must be removed after training. Evaluating such deletion is difficult because a lower membership score can reflect trajectory removal, residual memorization visible to another attack, or policy collapse that destroys useful behavior. We introduce Trajectory-level memOrization and Unlearning in offline RL (TOUR), a benchmark that combines trajectory-level partitioning, matched non-member controls, retraining references, retained-performance anchors, and multi-attack privacy auditing. Across D4RL locomotion experiments and an exploratory AntMaze extension, TOUR shows that common deletion baselines have environment-dependent privacy-utility behavior. Retraining and fine-tuning often provide stronger retained-utility references than uniform GA+Refit, while TrajDeleter remains a useful comparator but is not uniformly stronger under the same audit. Reference-model, threshold, deviation, equivalence, action-error, representation-based, and query-limited attacks further show that a single likelihood-based membership score can overstate deletion quality. In the evaluated settings, conclusions about offline RL unlearning are therefore not stable under single-score auditing. They depend on matched non-member construction, retraining-relative calibration, attack family, retained utility, and explicit scope for diagnostic architecture or component-level evidence.

强化学习数据隐私遗忘机制评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。