arXiv:2608.20784cs.ROcs.LG2026-08

让机器人模型删除特定示范,同时确保行为不变且无法被检测到。

Rethinking Demonstration Unlearning in Imitation Learning for Robotics

论文配图:Rethinking Demonstration Unlearning in Imitation Learning for Robotics
图 1 · 摘自论文原文
  • 设计双维度审计:行为一致性与证据可检测性,评估编辑效果。
  • 在真实机器人上实现18/20次成功,修复后任务表现接近重新训练。
  • 提出新验证方法,防止虚假“删除”伪装成有效编辑。

机器人模仿学习依赖人类示范,但部分示范可能后期需删除。直接重新训练成本高,因此需更低成本的策略来修改已训练模型。现有机器遗忘指标(如遗忘损失或单次成员攻击)无法准确衡量闭环执行中策略是否真正移除了示范影响。为此,本文提出一种重训校准的审计方法,从两个维度评估示范删除效果:行为维度衡量编辑后策略与重新训练结果在相同状态下的动作差异,通过独立重训构建基线;证据维度采用逐示范成员攻击,报告排名与绝对成员损失值,避免仅看排名导致的误判。通过置信区间检验结合两维结果,建立联合重训一致性假设。在三个真实机器人策略与两个仿真环境的五组预注册实验中,行为与证据可解耦——某些编辑能恢复任务性能但残留证据,或消除证据却偏离重训行为。在ACT机械臂实验中,定向编辑使盲评成功率恢复至20次中的18次。

原文摘要 · Abstract (English)

Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. Metrics inherited from machine unlearning, such as forgetting loss or a single membership attack, do not establish what an edit removed from a policy acting in closed loop. We therefore introduce a retrain-calibrated audit that reads demonstration unlearning along two axes: behavior, whether the edited policy acts like one retrained without the removed demonstrations, and evidence, whether an auditor can still detect it was trained on them. The behavior axis measures action divergence to that retrain at matched states, calibrated by a floor built from independent retrains, so a policy at the floor is as close to a retrain as retrains are to each other. The evidence axis applies a per-demonstration membership attack against a retrain null, reporting both its rank and its absolute member-loss level, since rank alone accepts operators that inflate member losses past the null. A conformal test then combines both axes into one hypothesis of joint retrain consistency, against a fleet of independent retrains large enough to reject at conventional significance. Across five preregistered conditions on three real-robot policy classes and two simulation suites, the axes dissociate in both directions on one checkpoint, as an edit may repair task behavior while leaving evidence unchanged, or reduce evidence while moving behavior away from retraining. On the ACT arm, a redirect edit restores blind-scored robot success to 18 of 20 trials.

模仿学习模型编辑机器人数据删除

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。