arXiv:2504.19854cs.ROcs.AI2025-04被引 108

3B参数小模型NORA实现高效机器人任务执行

NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks

  • 基于Qwen-2.5-VL-3B构建,融合视觉语义理解与动作定位
  • 在97万条真实机器人示范上训练,推理效率显著提升
  • 适合实时机器人系统,兼顾性能与计算开销

现有视觉-语言-动作(VLA)模型在零样本场景中表现优异,具备出色的任务执行与推理能力。然而,视觉编码限制常导致物体抓取等任务失败,且模型规模庞大(通常超过7B参数),带来高计算开销,难以满足实时机器人环境对速度与效率的要求。为此,本文提出NORA,一个3B参数的小型开源通用视觉语言动作模型。NORA以Qwen-2.5-VL-3B为骨干,利用其强大的视觉-语义理解能力提升视觉推理与动作接地性能,并在97万条真实世界机器人示范数据上训练,结合FAST+分词器实现高效动作序列生成。实验表明,NORA在保持高性能的同时显著降低计算开销,优于现有大型VLA模型,更适用于实时机器人自主任务。

原文摘要 · Abstract (English)

Existing Visual-Language-Action (VLA) models have shown promising performance in zero-shot scenarios, demonstrating impressive task execution and reasoning capabilities. However, a significant challenge arises from the limitations of visual encoding, which can result in failures during tasks such as object grasping. Moreover, these models typically suffer from high computational overhead due to their large sizes, often exceeding 7B parameters. While these models excel in reasoning and task planning, the substantial computational overhead they incur makes them impractical for real-time robotic environments, where speed and efficiency are paramount. To address the limitations of existing VLA models, we propose NORA, a 3B-parameter model designed to reduce computational overhead while maintaining strong task performance. NORA adopts the Qwen-2.5-VL-3B multimodal model as its backbone, leveraging its superior visual-semantic understanding to enhance visual reasoning and action grounding. Additionally, our \model{} is trained on 970k real-world robot demonstrations and equipped with the FAST+ tokenizer for efficient action sequence generation. Experimental results demonstrate that NORA outperforms existing large-scale VLA models, achieving better task performance with significantly reduced computational overhead, making it a more practical solution for real-time robotic autonomy.

机器人小模型视觉语言动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。