解决视觉模仿策略在相似物体前出错的问题,提升复杂任务下的目标选择鲁棒性。
What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

- 通过分阶段分析定位失败发生在抓取与放置环节
- 引入干扰物后性能下降30%以上,且敏感性依赖视觉相似性和操作阶段
- 提出注意力正则化与视觉提示等方法,显著提升仿真与真实机器人表现
视觉运动模仿策略在分布内条件下表现优异,但在引入视觉相似的物体或容器时会失效。本文将此问题归结为条件视觉定位:成功控制所需的视觉目标随操作阶段和任务状态而变化。基于动作分块变换器(ACT),我们系统地引入具有可控颜色和形状相似性的干扰物,定位到失败主要发生在抓取与放置阶段。研究发现,干扰敏感性同时依赖于视觉相似类型与操作阶段。据此诊断,我们评估了干扰增强、阶段依赖注意力正则化及基于外观的视觉提示等互补干预手段,在保持空间控制信息的前提下显著提升目标选择鲁棒性。这些方法在仿真与物理UR3e机器人上均有效。进一步在预训练视觉-语言-动作策略上验证了相同失效模式,该策略需根据医疗工具状态决定正确目的地。结果表明,即使底层操作技能完好,视觉干扰仍会导致错误对象或目的地选择;而显式改进目标选择可大幅恢复不同视觉运动策略中的性能。
原文摘要 · Abstract (English)
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。