通过分析动作漂移,精准挑选修复机器人模仿的薄弱数据。
It's Not Just More Demos: Counterfactual Action Sensitivity Coverage for Data-Efficient Robust Robot Imitation

- 基于动作漂移检测敏感性,筛选针对性修复数据
- 仅用20-30个样本即超越随机选择,接近大样本效果
- 适合数据效率与鲁棒性并重的机器人学习研究者
视觉运动模仿学习在操作任务中表现良好,但对光照、干扰或颜色变化等视觉扰动仍显脆弱。尽管增加数据多样性可提升鲁棒性,但难以判断哪些新增示范对特定策略最有价值。本文提出反事实干扰行为克隆(CFNBC),一种离线数据选择框架,用于针对性修复鲁棒性。从干净示范训练的基准策略出发,生成保持专家动作的成对干净与干扰观测,测量动作漂移——即在不应改变行为的干扰下策略预测动作的变化。该信号提供策略特异性的敏感度指标,可在不依赖在线执行或成功标签的前提下,从候选池中选出紧凑且响应多样的修复集。在MuJoCo双臂立方体搬运和SimplerEnv立方体堆叠任务中,动作漂移与干扰导致失败高度相关;使用仅20–30个选定样本的响应引导修复,显著优于同预算随机选择,并接近大规模随机修复的性能。结果支持以数据为中心的鲁棒性修复观点:最有价值的数据并非最多样或最难,而是覆盖当前策略脆弱响应模式的样本。
原文摘要 · Abstract (English)
Visuomotor imitation learning has demonstrated success for manipulation tasks. However, the trained policies remain brittle to visual `nuisances', with even minor task-preserving variations such as lighting, distractions or changes in colour result in heavy degradation of the trained policy's performance. While increasing data diversity can improve robustness, it is unclear which additional demonstrations are informative for a particular trained policy. We propose Counterfactual Nuisance Behaviour Cloning (CFNBC), an offline data-selection framework for targeted robustness repair. Starting from a nominal policy trained on `clean' demonstrations, CFNBC generates paired clean and nuisance observations that preserve the expert action, then measures \emph{action drift}: the change in the policy's predicted action under a nuisance that should not alter the desired behaviour. This provides a policy-specific sensitivity signal for selecting a compact, response-diverse repair set from a larger candidate pool, without requiring rollout success labels or online policy execution. We show in MuJoCo bimanual cube transfer and SimplerEnv cube stacking that action drift correlates with nuisance-induced failure, and that response-guided repair with only $20$--$30$ selected candidates substantially outperforms matched-budget random selection while approaching the performance of much larger random repair budgets. These results support a data-centric view of robustness repair: the most useful data are not necessarily the most numerous, visually diverse, or obviously difficult, but the examples that cover fragile response modes of the current policy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。