arXiv:2508.03660cs.LG2025-08被引 1

用少量参数微调让机器人政策快速适配新身体结构。

Efficient Morphology-Aware Policy Transfer to New Embodiments

  • 结合预训练与参数高效微调技术,减少适配新机器人的参数量。
  • 仅调整不到1%参数就能显著提升零样本性能。
  • 适合需要快速部署、数据采集成本高的机器人场景。

形态感知策略学习通过整合多智能体数据提升策略样本效率,能有效泛化于动态、运动学及肢体配置变化。然而,其在部署时的零样本性能仍逊于针对具体形态的端到端微调。这在机器人应用中带来挑战,因进一步数据收集进行端到端微调成本高昂。本文研究将形态感知预训练与参数高效微调(PEFT)结合,以降低将策略适配至目标形态所需的学习参数量。比较了直接微调部分权重、可学习适配器和前缀微调等在线微调方法。结果表明,结合预训练的PEFT技术显著减少达到良好性能所需的样本数;即使仅调整少于总参数的1%,也能优于基线预训练策略的零样本表现。

原文摘要 · Abstract (English)

Morphology-aware policy learning is a means of enhancing policy sample efficiency by aggregating data from multiple agents. These types of policies have previously been shown to help generalize over dynamic, kinematic, and limb configuration variations between agent morphologies. Unfortunately, these policies still have sub-optimal zero-shot performance compared to end-to-end finetuning on morphologies at deployment. This limitation has ramifications in practical applications such as robotics because further data collection to perform end-to-end finetuning can be computationally expensive. In this work, we investigate combining morphology-aware pretraining with parameter efficient finetuning (PEFT) techniques to help reduce the learnable parameters necessary to specialize a morphology-aware policy to a target embodiment. We compare directly tuning sub-sets of model weights, input learnable adapters, and prefix tuning techniques for online finetuning. Our analysis reveals that PEFT techniques in conjunction with policy pre-training generally help reduce the number of samples to necessary to improve a policy compared to training models end-to-end from scratch. We further find that tuning as few as less than 1% of total parameters will improve policy performance compared the zero-shot performance of the base pretrained a policy.

机器人策略迁移参数高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。