arXiv:2601.19514cs.RO2026-01被引 3

通过视觉对齐提升机器人抓取策略在域外场景下的泛化能力

PALM: Enhanced Generalizability for Local Visuomotor Policies via Perception Alignment

  • 将操作策略分解为全局与局部两部分,聚焦局部动作的视觉不变性
  • 在仿真中域外性能下降仅8%,真实世界下降24%,远优于基线
  • 无需额外数据或模型改动,适用于多种物理环境和机械臂迁移

基于图像的行为克隆在域外泛化方面仍具挑战。现有方法分别处理工作空间变化、视角差异和跨体态迁移等问题,但通常独立开发且依赖复杂流程。我们提出PALM(局部操作的感知对齐),利用域外(OOD)与演示域间局部动作分布的不变性,同时应对多种域外偏差,无需额外输入模态、模型修改或数据收集。PALM将操作策略模块化为粗粒度全局组件与细粒度局部策略。通过强制局部视觉聚焦和一致本体感觉表征,减小局部策略层面的域内与域外输入差异,使策略能在域外条件下检索出不变的局部动作。实验表明,相比基线在仿真中性能下降45%、真实世界下降77%,PALM仅分别下降8%和24%。

原文摘要 · Abstract (English)

Generalizing beyond the training domain in image-based behavior cloning remains challenging. Existing methods address individual axes of generalization, workspace shifts, viewpoint changes, and cross-embodiment transfer, yet they are typically developed in isolation and often rely on complex pipelines. We introduce PALM (Perception Alignment for Local Manipulation), which leverages the invariance of local action distributions between out-of-distribution (OOD) and demonstrated domains to address these OOD shifts concurrently, without additional input modalities, model changes, or data collection. PALM modularizes the manipulation policy into coarse global components and a local policy for fine-grained actions. We reduce the discrepancy between in-domain and OOD inputs at the local policy level by enforcing local visual focus and consistent proprioceptive representation, allowing the policy to retrieve invariant local actions under OOD conditions. Experiments show that PALM limits OOD performance drops to 8% in simulation and 24% in the real world, compared to 45% and 77% for baselines.

机器人控制视觉对齐泛化能力行为克隆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。