arXiv:2502.00543cs.ROcs.CV2025-02被引 7

用一小时数据训练多任务机器人越野模型,实现高效路径预测与控制。

VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility

  • 设计新型非自回归掩码建模,同时预测姿态、动作和地形
  • 仅用1小时真实数据训练,支持正向/逆向运动学建模等多任务
  • 适合数据稀缺场景下机器人自主导航系统研发者使用

复杂的学习架构如Transformer为机器人理解复杂车辆-地形动力学交互提供了新可能。尽管自然语言处理(NLP)和计算机视觉(CV)可利用互联网规模数据训练Transformer,但真实机器人在崎岖垂直地形上获取移动数据极为困难。此外,专为文本和图像设计的训练技术难以直接应用于机器人移动任务。本文提出VertiFormer,一种新型数据高效的多任务Transformer模型,仅需一小时真实数据即可在极端陡峭的越野地形上实现机器人机动性建模。VertiFormer采用新型可学习掩码建模与下一词预测机制,同时预测下一姿态、动作及地形块,支持前向与逆向运动学建模等多种任务。其非自回归设计缓解了自回归模型的计算瓶颈与误差传播问题。统一模态表示增强了对多样化时序映射与状态表征的学习能力,结合多目标函数进一步提升模型泛化性。实验表明,该模型能在有限数据下有效应用Transformer于越野机器人机动性,并可在物理机器人上实现实时多任务支持。

原文摘要 · Abstract (English)

Sophisticated learning architectures, e.g., Transformers, present a unique opportunity for robots to understand complex vehicle-terrain kinodynamic interactions for off-road mobility. While internet-scale data are available for Natural Language Processing (NLP) and Computer Vision (CV) tasks to train Transformers, real-world mobility data are difficult to acquire with physical robots navigating off-road terrain. Furthermore, training techniques specifically designed to process text and image data in NLP and CV may not apply to robot mobility. In this paper, we propose VertiFormer, a novel data-efficient multi-task Transformer model trained with only one hour of data to address such challenges of applying Transformer architectures for robot mobility on extremely rugged, vertically challenging, off-road terrain. Specifically, VertiFormer employs a new learnable masked modeling and next token prediction paradigm to predict the next pose, action, and terrain patch to enable a variety of off-road mobility tasks simultaneously, e.g., forward and inverse kinodynamics modeling. The non-autoregressive design mitigates computational bottlenecks and error propagation associated with autoregressive models. VertiFormer's unified modality representation also enhances learning of diverse temporal mappings and state representations, which, combined with multiple objective functions, further improves model generalization. Our experiments offer insights into effectively utilizing Transformers for off-road robot mobility with limited data and demonstrate our efficiently trained Transformer can facilitate multiple off-road mobility tasks onboard a physical mobile robot.

机器人Transformer多任务数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。