提出轻量级机器人推理新策略,提速3倍且性能更强
Training Strategies for Efficient Embodied Reasoning
- 分离验证推理提升性能的三种机制:表征学习、课程学习、表达力增强
- 在LIBERO-90上达到当前最优表现,推理速度提升3倍
- 适合追求高效视觉语言动作模型的机器人研发者
机器人链式思维推理(CoT)——即模型在执行动作前预测有用中间表示——能有效提升机器人策略的泛化能力和性能,尤其在视觉-语言-动作模型(VLAs)中。尽管已有研究证明其有效性,但该方法仍存在需专用推理数据、推理速度慢等核心问题。为此,本文提出三种可能机制:(1) 更优的表征学习,(2) 改进的学习课程化,(3) 增强的表达能力,并设计简单变体来分别验证。结果表明,生成推理过程确实能提升VLA的表征质量,而关注这些推理有助于更有效地用于动作预测。基于此,提出两种轻量级替代方案,在不牺牲性能的前提下实现显著加速。所提方法在非推理策略上取得显著性能提升,在LIBERO-90基准上达到当前最佳水平,并实现相比标准推理3倍的推理速度提升。
原文摘要 · Abstract (English)
Robot chain-of-thought reasoning (CoT) -- wherein a model predicts helpful intermediate representations before choosing actions -- provides an effective method for improving the generalization and performance of robot policies, especially vision-language-action models (VLAs). While such approaches have been shown to improve performance and generalization, they suffer from core limitations, like needing specialized robot reasoning data and slow inference speeds. To design new robot reasoning approaches that address these issues, a more complete characterization of why reasoning helps policy performance is critical. We hypothesize several mechanisms by which robot reasoning improves policies -- (1) better representation learning, (2) improved learning curricularization, and (3) increased expressivity -- then devise simple variants of robot CoT reasoning to isolate and test each one. We find that learning to generate reasonings does lead to better VLA representations, while attending to the reasonings aids in actually leveraging these features for improved action prediction. Our results provide us with a better understanding of why CoT reasoning helps VLAs, which we use to introduce two simple and lightweight alternative recipes for robot reasoning. Our proposed approaches achieve significant performance gains over non-reasoning policies, state-of-the-art results on the LIBERO-90 benchmark, and a 3x inference speedup compared to standard robot reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。