arXiv:2509.23220cs.RO2025-09

通过追踪关键局部区域,提升机器人在复杂场景下的模仿学习鲁棒性。

GLUE: Global-Local Unified Encoding for Imitation Learning via Key-Patch Tracking

  • 用文本引导追踪关键局部块,融合全局与局部特征
  • 仿真和真实环境分别提升17.6%和36.3%性能
  • 特别适合处理遮挡、杂乱等分布外场景

近年来,视觉表征学习在机器人模仿学习中受到广泛关注。然而,在由杂乱和遮挡导致的分布外(OOD)场景中,全局视觉表征的注意力可能被稀释或干扰,导致策略性能下降。局部表征对任务相关物体的不变性提供了解决方案。通过高效利用这些局部表征,可将训练与测试数据映射到更相似的特征空间,缓解协变量偏移问题。为此,我们提出GLUE——一种基于关键块追踪的全局-局部统一编码框架。GLUE采用文本引导机制选择并追踪关键块作为重要局部表征,设计了一种新型融合框架:全局块特征查询局部块以提炼关键信息,生成与全局上下文异质性低的细粒度局部特征。该融合表示引导机器人视觉注意力聚焦于任务相关物体,同时保留精确的全局上下文,使训练与测试分布对齐至相似且任务相关的特征空间,最终提升模仿学习策略的鲁棒性。实验表明,GLUE在多种任务的仿真与真实世界设置中均表现优异,相比最强基线,仿真提升17.6%,真实环境提升36.3%,真实世界泛化设置提升58.3%。

原文摘要 · Abstract (English)

In recent years, visual representation learning has gained widespread attention in robotic imitation learning. However, in complex Out-of-Distribution(OOD) settings characterized by clutter and occlusion, the attention of global visual representations can be diluted or interfered, leading to degraded policy performance. The invariance of local representations for task-relevant objects offers a solution. By efficiently utilizing these local representations, training and testing data can be mapped to a more similar feature space, thereby mitigating the covariate shift problem. Accordingly, we propose GLUE, a global-local unified encoding framework for imitation learning based on key-patch tracking. GLUE selects and tracks key-patches as critical local representations by employing a text-guided mechanism. It features a novel fusion framework where global patch features query local patches to distill essential information, yielding fine-grained local features with low heterogeneity relative to the global context. This fused representation steers the robot's visual attention toward task-relevant objects and preserves precise global context, which together align the training and testing distributions into a similar and task-informative feature space, ultimately enhancing the robustness of the imitation learning policy. Experiments demonstrate that GLUE achieves strong performance across diverse tasks in both simulation and real-world settings, outperforming the strongest baseline by 17.6% in simulation, 36.3% in real-world environments, and 58.3% on real-world generalization settings. The project website of GLUE is available at https://GLUE666.github.io/.

模仿学习视觉表征机器人关键块追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。