arXiv:2607.10172cs.ROcs.LG2026-07中稿 · the International …

LoRA微调可让大模型在工业机器人上高效运行,32秩时性能达最优且显存降低70%

On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation

论文配图:On the Efficiency of LoRA Fine-Tuning for Vision-Language-Action Models in Industrial Robotic Manipulation
图 1 · 摘自论文原文
  • 用32秩LoRA对视觉-语言-动作模型进行全编码器微调
  • 性能与全参数微调相当,显存峰值从36.2降至10.8 GiB
  • 适合资源受限的工业机器人部署,尤其需视觉语义协同适应

将百亿参数视觉-语言-动作(VLA)模型部署于工业硬件需微调以弥合具身差距。全参数微调(FFT)虽具最大可塑性,但需数据中心级GPU。本文系统研究了针对π₀(基于流匹配的VLA)的低秩适配(LoRA),在四类精密装配任务中使用UR5e机械臂评估。在不同LoRA秩(r=8至256)、分配策略及组件冻结消融实验中,未发现FFT在统计学上优于某些LoRA配置。性能在r=32时饱和,且对视觉-语言模型主干与动作专家均匀分配足够。冻结视觉-语言模型或仅对视觉编码器启用LoRA会显著降低性能,表明具身适应需同时具备语义与视觉可塑性。结果表明,采用r=32的完整视觉编码器微调方案是实用选择,静态峰值显存从36.2降至10.8 GiB(不含参数、优化器状态及激活内存),性能无明显损失。

原文摘要 · Abstract (English)

Deploying billion-parameter Vision-Language-Action (VLA) models on industrial hardware requires fine-tuning to bridge the embodiment gap. Full Fine-Tuning (FFT) provides maximal plasticity but requires data centre-grade GPUs. We present a systematic study of Low-Rank Adaptation (LoRA) for $π_0$, a flow-matching VLA, evaluated on four precision assembly tasks with a UR5e robotic manipulator. Across a sweep of LoRA ranks (r=8 to 256), allocation strategies, and component-freezing ablations, we find no statistically significant advantage of FFT over certain LoRA configurations. Performance saturates at r=32, and uniform allocation across the Vision-Language-Model (VLM) backbone and action expert proves sufficient. Freezing the VLM or restricting the vision encoder to LoRA significantly degrades performance, indicating that embodiment adaptation requires both semantic and visual plasticity. These results suggest that LoRA at r=32 with full vision encoder fine-tuning is a practical approach, reducing static peak VRAM from 36.2 to 10.8 GiB (parameters and optimizer states, activation memory excluded) without detectable performance loss.

LoRA机器人控制视觉语言模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。