轻量微调让视觉语言动作模型更抗视角变化
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
- 用轻量参数重校视觉表征,修复视角对齐问题
- 仅4千参数使视点准确率从48.5%提升至87.1%
- 适合想低成本提升模型鲁棒性的研究者
视觉-语言-动作(VLA)模型在分布内表现优异,但在新相机视角和视觉扰动下性能急剧下降。我们发现这种脆弱性主要源于空间建模的错位,而非物理建模问题。为此,提出一种一次性适应框架,通过轻量可学习更新重新校准视觉表征。首个方法特征令牌调制(FTM)对视觉令牌施加全局仿射变换,在仅使用4K参数的情况下,将Libero数据集上的视点准确率从48.5%提升至87.1%。在此基础上,特征线性适配(FLA)引入低秩更新至ViT编码器,以470万参数实现90.8%成功率,达到LoRA级微调效果但成本远低于后者。结果表明预训练VLA模型具备大量未开发的鲁棒性,且针对性、极小规模的视觉适应足以恢复视角泛化能力。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose a one-shot adaptation framework that recalibrates visual representations through lightweight, learnable updates. Our first method, Feature Token Modulation (FTM), applies a global affine transformation to visual tokens and improves Libero viewpoint accuracy from 48.5% to 87.1% with only 4K parameters. Building on this, Feature Linear Adaptation (FLA) introduces low-rank updates to the ViT encoder, achieving 90.8% success with 4.7M parameters -- matching LoRA-scale finetuning at far lower cost. Together, these results reveal substantial untapped robustness in pretrained VLA models and demonstrate that targeted, minimal visual adaptation is sufficient to restore viewpoint generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。