arXiv:2509.11839cs.ROcs.CV2025-09被引 6

用轮式人形数据提升双足人形操作能力,仅需10分钟实操即可落地。

TrajBooster: Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning

  • 以末端轨迹为通用接口,跨形态迁移动作策略
  • 在Unitree G1上实现蹲姿、跨高度操作等复杂任务
  • 仅需10分钟真实数据即可完成部署,适合双足机器人快速适配

近期视觉-语言-动作模型虽具跨体态泛化潜力,但在缺乏高质量示范时难以快速适配新机器人动作空间,尤其对双足人形尤为困难。我们提出TrajBooster,一种跨体态框架,利用丰富的轮式人形数据增强双足人形VLA性能。核心思路是将末端执行器轨迹作为与形态无关的接口:(i) 从真实轮式人形中提取6D双臂末端轨迹;(ii) 在仿真中将轨迹重定向至Unitree G1,通过启发式增强的协同在线DAgger训练全身控制器,将低维轨迹参考转化为高维可行动作;(iii) 构建异构三元组,耦合源视觉/语言输入与目标人形兼容动作,用于后预训练VLA,随后仅需10分钟远程操控数据采集即可在目标人形域部署。在Unitree G1上,该策略成功实现桌面上下及跨高度家务任务,支持蹲姿、协调全身运动,显著提升鲁棒性与泛化能力。结果表明,TrajBooster能高效利用现有轮式人形数据强化双足人形VLA表现,降低对昂贵同体态数据依赖,增强动作空间理解与零样本技能迁移能力。

原文摘要 · Abstract (English)

Recent Vision-Language-Action models show potential to generalize across embodiments but struggle to quickly align with a new robot's action space when high-quality demonstrations are scarce, especially for bipedal humanoids. We present TrajBooster, a cross-embodiment framework that leverages abundant wheeled-humanoid data to boost bipedal VLA. Our key idea is to use end-effector trajectories as a morphology-agnostic interface. TrajBooster (i) extracts 6D dual-arm end-effector trajectories from real-world wheeled humanoids, (ii) retargets them in simulation to Unitree G1 with a whole-body controller trained via a heuristic-enhanced harmonized online DAgger to lift low-dimensional trajectory references into feasible high-dimensional whole-body actions, and (iii) forms heterogeneous triplets that couple source vision/language with target humanoid-compatible actions to post-pre-train a VLA, followed by only 10 minutes of teleoperation data collection on the target humanoid domain. Deployed on Unitree G1, our policy achieves beyond-tabletop household tasks, enabling squatting, cross-height manipulation, and coordinated whole-body motion with markedly improved robustness and generalization. Results show that TrajBooster allows existing wheeled-humanoid data to efficiently strengthen bipedal humanoid VLA performance, reducing reliance on costly same-embodiment data while enhancing action space understanding and zero-shot skill transfer capabilities. For more details, For more details, please refer to our \href{https://jiachengliu3.github.io/TrajBooster/}.

人形机器人动作迁移视觉语言动作轨迹学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。