arXiv:2512.11921cs.ROcs.AI2025-12被引 4

用低资源方法让大模型在便宜机器人上跑起来,实现自然语言控制。

Towards Accessible Physical AI: LoRA-Based Fine-Tuning of VLA Models for Real-World Robot Control

  • 用LoRA和量化技术微调31亿参数大模型,只用8GB显存
  • 200次示范数据就能让机器人完成按按钮任务,性能稳定
  • 适合想低成本部署智能机器人的研究者和开发者

视觉-语言-动作(VLA)模型在机器人操作中展现出强大能力,可直接通过视觉观察执行自然语言指令。然而,由于计算限制及对新机器人形态的高效适配需求,将大规模VLA模型部署在廉价机器人平台上仍具挑战。本文提出一种资源高效的微调方法与真实世界部署分析,将31亿参数的VLA模型适配至低成本机器人系统。采用低秩适应(LoRA)与量化技术,使模型可在仅8GB显存的消费级GPU上运行。针对有限示范数据下的新机器人形态适配问题,研究了冻结与非冻结视觉编码器的权衡。在SO101机械臂上进行按按钮任务的真实部署测试,基于200个示范回合训练,验证了方法在保持计算效率的同时实现有效操作性能。结果表明,通过合理微调策略,大型VLA模型可成功部署于经济型机器人平台,使先进操作能力不再局限于昂贵科研设备。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in robotic manipulation,enabling robots to execute natural language commands through end-to-end learning from visual observations.However, deploying large-scale VLA models on affordable robotic platforms remains challenging due to computational constraints and the need for efficient adaptation to new robot embodiments. This paper presents an efficient fine-tuning methodology and real-world deployment analysis for adapting VLA models to low-cost robotic manipulation systems.We propose a resource-efficient fine-tuning strategy using Low-Rank Adaptation (LoRA) and quantization techniques that enable multi-billion parameter VLA models ( 3.1B parameters) to run on consumer-grade GPUs with 8GB VRAM. Our methodology addresses the critical challenge of adapting pre-trained VLA models to new robot embodiments with limited demonstration data, focusing on the trade-offs between frozen and unfrozen vision encoders. Through real-world deployment on the SO101 robotic arm for a button-pressing manipulation task, we demonstrate that our approach achieves effective manipulation performance while maintaining computational efficiency. We provide detailed analysis of deployment challenges, failure modes, and the relationship between training data quantity and real-world performance,trained on 200 demonstration episodes. Our results show that with proper fine-tuning methodology, VLA models can be successfully deployed on affordable robotic platforms,making advanced manipulation capabilities accessible beyond expensive research robots.

机器人控制LoRA微调低成本部署VLA模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。