arXiv:2603.14554cs.RO2026-03

通过修正价值函数校准问题,实现四足机器人零样本跨形态迁移

MorFiC: Fixing Value Miscalibration for Zero-Shot Quadruped Transfer

  • 用形态隐变量乘性条件化批评者,解决多形态间价值目标冲突
  • 单机器人训练2小时即实现7种机器人的零样本迁移,最高速度1.98米/秒
  • 适合追求高效低资源跨形态强化学习的机器人研发团队

在不同形态的四足机器人之间泛化已学习的运动策略仍是挑战。在单一机器人上训练的策略往往在具有不同质量分布、运动学特性、关节限制或驱动约束的机器人上部署失败,需重新训练。现有方法主要依赖在真实或生成的多种形态上进行训练,或使用大型模型以获得可迁移策略,但均带来高存储与计算开销。我们指出,行为-评价学习中的一个关键失败模式是:共享的价值函数会平均不兼容的价值目标,导致优势值校准错误。为此提出MorFiC,一种基于形态隐变量乘性条件化的强化学习方法,可在模拟中对单一源机器人进行形态随机化训练。该方法在无任何微调的情况下成功实现零样本迁移至7种不同机器人,其中AlienGo上达到1.98米/秒的速度,显著优于加性条件的PPO基线(<0.65米/秒)。整个训练仅耗时2小时,远低于其他方法数十至数百小时。通过解释方差、优势符号翻转率和策略梯度余弦三项指标验证了批评者校准的有效性。最终实现在Unitree Go1与Go2上的零样本部署。

原文摘要 · Abstract (English)

Generalizing learned locomotion policies across quadrupedal robots with different morphologies remains a challenge. Policies trained on a single robot often fail when deployed on embodiments with different mass distributions, kinematics, joint limits, or actuation constraints, forcing per-robot retraining. Prior works have approached this primarily either by scaling via training across real or generated embodiments or using large architectures producing transferable policies which are both storage and compute heavy. We argue that both of these approaches circumvent a key failure mode in actor-critic learning: a shared value function tends to average incompatible value targets across embodiments, yielding miscalibrated advantages and fixing that helps transfer and reduce the compute cost. We present MorFiC, a reinforcement learning approach for zero-shot cross-morphology locomotion which fixes this issue using multiplicative critic conditioning on a morphology latent. Trained with a single source robot with morphology randomization in simulation, MorFiC achieves zero-shot transfer to total seven robots and reaching forward velocity competitive or surpassing scaling-based and morphology-conditioned PPO baselines for examples 1.98 m/s on AlienGo, where as additive PPO baselines remain below 0.65 m/s. MorFiC achieves this using single training robot within 2 hours of training time compared to tens or hundreds of hours required for scaling baselines training on multiple embodiments. We diagnose critic calibration via three metric: Explained variance, advantage sign-flip rate and policy gradient cosine, confirming our critic conditioning as determinant for policy update direction. Finally, we demonstrate zero-shot deployment on Unitree Go1 and Go2 robots without fine-tuning.

强化学习四足机器人零样本迁移价值校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。