arXiv:2506.06196cs.RO2025-06被引 4

用空间化中间表征提升机器人灵巧操作的泛化能力

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

  • 设计物体中心、姿态感知、深度感知的中间表征,指导策略学习
  • 新混合专家架构使成功率比语言基线高11%,比标准扩散策略高24%
  • 结合加权模仿学习可进一步提升10%性能,适合复杂抓取任务研究者

本文研究空间化辅助表征如何同时提供高层语义接地与直接可操作信息,以提升灵巧双臂操作任务中的策略学习性能与泛化能力。我们从物体中心性、姿态感知和深度感知三个维度评估中层表征,并通过监督学习训练专用编码器,将其输入扩散策略解决真实世界中的灵巧操作任务。提出一种新型混合专家策略架构,整合多个基于不同中层表征的专家模型,显著提升泛化能力。该方法在评估任务上平均成功率较语言基线高出11%,较标准扩散策略高出24%。此外,在加权模仿学习框架中利用中层表征作为动作监督信号,使策略对表征的遵循精度提升,带来额外10%的性能增益。结果表明,将机器人策略不仅锚定于广泛感知任务,更需结合细粒度、可操作的表征。

原文摘要 · Abstract (English)

In this work, we investigate how spatially grounded auxiliary representations can provide both broad, high-level grounding as well as direct, actionable information to improve policy learning performance and generalization for dexterous tasks. We study these mid-level representations across three critical dimensions: object-centricity, pose-awareness, and depth-awareness. We use these interpretable mid-level representations to train specialist encoders via supervised learning, then feed them as inputs to a diffusion policy to solve dexterous bimanual manipulation tasks in the real world. We propose a novel mixture-of-experts policy architecture that combines multiple specialized expert models, each trained on a distinct mid-level representation, to improve policy generalization. This method achieves an average success rate that is 11% higher than a language-grounded baseline and 24 percent higher than a standard diffusion policy baseline on our evaluation tasks. Furthermore, we find that leveraging mid-level representations as supervision signals for policy actions within a weighted imitation learning algorithm improves the precision with which the policy follows these representations, yielding an additional performance increase of 10%. Our findings highlight the importance of grounding robot policies not only with broad perceptual tasks but also with more granular, actionable representations. For further information and videos, please visit https://mid-level-moe.github.io.

机器人策略学习中层表征扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。