arXiv:2603.25661cs.ROcs.CV2026-03被引 7

用轻量方法提升视觉语言动作模型性能,训练更快更省资源。

Fast-dVLA: Accelerating Discrete Diffusion VLA to Real-Time Performance

  • 分离通用能力与任务适配目标,分两阶段训练
  • 仅用小规模任务集即可生成能力向量,提升模型泛化
  • 融合预训练参数后,性能媲美复杂方法且开销更低

本文针对预训练视觉语言动作(VLA)模型在标准监督微调(SFT)中难以有效提升性能且适应成本高的问题,提出一种新方法。现有先进微调方法虽能提升性能并减少收敛步数,但常因附加任务损失带来显著计算开销。为此,本文在参数空间中解耦辅助任务训练的两个目标:增强通用能力与拟合特定任务动作分布。仅需在小规模任务集上使用两种不同训练策略进行训练,其参数差异即构成由辅助任务提供的能力向量。将该向量与预训练参数融合,形成能力增强的元模型。进一步地,当标准SFT加入轻量正交正则化损失时,该融合模型性能可媲美辅助微调基线,同时计算开销大幅降低。实验表明,该方法在多种机器人任务中均表现优异。

原文摘要 · Abstract (English)

This paper proposes a novel approach to address the challenge that pretrained VLA models often fail to effectively improve performance and reduce adaptation costs during standard supervised finetuning (SFT). Some advanced finetuning methods with auxiliary training objectives can improve performance and reduce the number of convergence steps. However, they typically incur significant computational overhead due to the additional losses from auxiliary tasks. To simultaneously achieve the enhanced capabilities of auxiliary training with the simplicity of standard SFT, we decouple the two objectives of auxiliary task training within the parameter space, namely, enhancing general capabilities and fitting task-specific action distributions. To deliver this goal, we only need to train the model to converge on a small-scale task set using two distinct training strategies. The difference between the resulting model parameters can then be interpreted as capability vectors provided by auxiliary tasks. These vectors are then merged with pretrained parameters to form a capability-enhanced meta model. Moreover, when standard SFT is augmented with a lightweight orthogonal regularization loss, the merged model attains performance comparable to auxiliary finetuned baselines with reduced computational overhead. Experimental results demonstrate that this approach is highly effective across diverse robot tasks. Project page: https://chris1220313648.github.io/Fast-dVLA/

视觉语言动作模型微调轻量化机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。