HiRT让机器人在保持低算力的同时实现高速动态操作。
HiRT: Enhancing Robotic Control with Hierarchical Robot Transformers
- 分层设计:低频运行视觉语言模型捕捉稳定特征,高频执行实时控制。
- 静态任务下控制频率翻倍,成功率相当;动态任务成功率从48%提升至75%。
- 适合需要快速响应的机器人应用场景,如真实世界动态抓取。
大型视觉-语言-动作(VLA)模型凭借强大的预训练视觉-语言模型(VLM)后端,在机器人控制中展现出优异的泛化能力。然而,其成功伴随高计算成本与推理延迟,导致主要适用于准静态任务,难以应对需快速交互的动态任务。为此,本文提出分层机器人变换器(HiRT),实现灵活的频率与性能权衡。HiRT使VLM以低频运行以捕获时间不变特征,同时通过由缓慢更新特征引导的高频视觉策略实现实时交互。仿真与真实场景实验表明,相比基线方法有显著提升:静态任务下控制频率翻倍,成功率相当;在以往VLA模型难以处理的新颖真实世界动态操作任务中,成功率从48%提升至75%。
原文摘要 · Abstract (English)
Large Vision-Language-Action (VLA) models, leveraging powerful pre trained Vision-Language Models (VLMs) backends, have shown promise in robotic control due to their impressive generalization ability. However, the success comes at a cost. Their reliance on VLM backends with billions of parameters leads to high computational costs and inference latency, limiting the testing scenarios to mainly quasi-static tasks and hindering performance in dynamic tasks requiring rapid interactions. To address these limitations, this paper proposes HiRT, a Hierarchical Robot Transformer framework that enables flexible frequency and performance trade-off. HiRT keeps VLMs running at low frequencies to capture temporarily invariant features while enabling real-time interaction through a high-frequency vision-based policy guided by the slowly updated features. Experiment results in both simulation and real-world settings demonstrate significant improvements over baseline methods. Empirically, in static tasks, we double the control frequency and achieve comparable success rates. Additionally, on novel real-world dynamic ma nipulation tasks which are challenging for previous VLA models, HiRT improves the success rate from 48% to 75%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。