让机器人在遇到新干扰物时仍能准确预测动作结果,提升视觉规划可靠性。
Reimagination with Test-time Observation Interventions: Distractor-Robust World Model Predictions for Visual Model Predictive Control
- 通过检测预测中物理上不合理的变化,识别并移除视觉干扰物
- 修改观测后重新想象未来状态,再事后还原干扰物保持视觉一致
- 在新干扰物下任务成功率最高提升3倍,适合开放世界机器人控制
世界模型使机器人能够基于当前观测和计划动作“想象”未来视觉状态,已被广泛用于通用动态建模以促进机器人学习。然而,当遇到训练中罕见的新型视觉干扰物(如新物体或背景元素)时,这些模型仍易失效。具体表现为:新干扰物会破坏动作结果预测,导致依赖世界模型进行规划或动作验证的机器人出现下游失败。本文提出测试时观察干预方法ReOI(Reimagination with Observation Intervention),一种简单而有效的策略,使世界模型在存在未预期视觉干扰的开放世界场景中仍能生成可靠的动作结果预测。给定当前机器人观测,ReOI首先通过识别世界模型预测过程中出现物理上不合理的退化成分来检测视觉干扰物;随后修改当前观测以移除这些干扰物,使其更接近训练分布;最后,利用修正后的观测重新生成未来状态,并在事后将干扰物重新引入,以维持下游规划与验证所需的视觉一致性。我们在一系列机器人操作任务中验证了该方法在动作验证场景下的有效性,结果显示,ReOI对分布内与分布外的视觉干扰物均具有鲁棒性。特别地,在存在新型干扰物时,任务成功率最高提升3倍,显著优于不使用想象干预的世界模型预测方案。
原文摘要 · Abstract (English)
World models enable robots to "imagine" future observations given current observations and planned actions, and have been increasingly adopted as generalized dynamics models to facilitate robot learning. Despite their promise, these models remain brittle when encountering novel visual distractors such as objects and background elements rarely seen during training. Specifically, novel distractors can corrupt action outcome predictions, causing downstream failures when robots rely on the world model imaginations for planning or action verification. In this work, we propose Reimagination with Observation Intervention (ReOI), a simple yet effective test-time strategy that enables world models to predict more reliable action outcomes in open-world scenarios where novel and unanticipated visual distractors are inevitable. Given the current robot observation, ReOI first detects visual distractors by identifying which elements of the scene degrade in physically implausible ways during world model prediction. Then, it modifies the current observation to remove these distractors and bring the observation closer to the training distribution. Finally, ReOI "reimagines" future outcomes with the modified observation and reintroduces the distractors post-hoc to preserve visual consistency for downstream planning and verification. We validate our approach on a suite of robotic manipulation tasks in the context of action verification, where the verifier needs to select desired action plans based on predictions from a world model. Our results show that ReOI is robust to both in-distribution and out-of-distribution visual distractors. Notably, it improves task success rates by up to 3x in the presence of novel distractors, significantly outperforming action verification that relies on world model predictions without imagination interventions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。