arXiv:2512.01358cs.ROcs.LG2025-12被引 1

通过增强感知模态,让机器人政策跨体型更稳定地完成抓取任务。

Modality-Augmented Fine-Tuning of Foundation Robot Policies for Cross-Embodiment Manipulation on GR1 and G1

论文配图:Modality-Augmented Fine-Tuning of Foundation Robot Policies for Cross-Embodiment Manipulation on GR1 and G1
图 1 · 摘自论文原文
  • 在现有数据上添加接触信号和深度图,提升模型对环境的感知能力。
  • 在GR1上成功率从51%提升至63%,在G1上抓苹果任务达94%。
  • 适合研究多模态融合与机器人跨平台迁移的团队参考。

本文提出一种模态增强的微调框架,用于将基础机器人策略适配到不同类人机器人形态。在两个场景中验证:(i) GR1形态,利用公开数据集并引入后处理模态,包括二值接触信号和ZoeDepth生成的度量深度;(ii) Unitree G1形态,贡献了一个新多模态数据集,包含cuRobo运动规划、逆运动学和真实接触力测量。实验表明,模态增强能持续提升不同形态下的策略性能。在GR1上,加入接触状态线索和RGB-D融合后,线上成功率从51%提升至63%。在G1的“抓苹果放入碗”任务中,接触增强模型达到94%成功率,显著优于标准微调的48%和零样本迁移的0%。结果表明,轻量级后处理可有效强化GR1策略,而高质量多模态数据对可靠迁移至Unitree G1至关重要。本工作建立了一条统一的数据驱动路径,通过针对性模态设计与多模态微调扩展基础机器人策略。

原文摘要 · Abstract (English)

This paper presents a modality-augmented fine-tuning framework designed to adapt foundation robot policies to diverse humanoid embodiments. We validate our approach across two distinct settings: (i) the GR1 embodiment, utilizing public datasets where we introduce post-processed modalities, including binary contact signals and ZoeDepth-generated metric depth; and (ii) the Unitree G1 embodiment, for which we contribute a novel multi-modal dataset incorporating cuRobo motion planning, inverse kinematics, and ground-truth contact-force measurements. Our experiments demonstrate that modality augmentation consistently enhances policy performance across different embodiments. Specifically, for the GR1, integrating contact-state cues and RGB-D fusion improves online success rates from 51% to 63%. Furthermore, in the G1 "Pick Apple to Bowl" task, our contact-augmented model achieves a success rate of 94%, significantly outperforming the 48% achieved by standard fine-tuning and the 0% baseline of zero-shot transfer. These results highlight that lightweight post-processing effectively strengthens policies for GR1, while high-quality multi-modal data is crucial for reliable transfer to the Unitree G1. Consequently, this work establishes a unified, data-centric pathway for extending foundation robot policies through targeted modality design and multi-modal fine-tuning.

机器人策略多模态跨体型微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。