arXiv:2606.31382cs.RO2026-06被引 1

发现视觉语言动作模型剪枝后性能下降的根源,提出按模块精准剪枝新方法。

Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation

论文配图:Revisiting Parameter Redundancy in Vision-Language-Action Models: Insights from VLM-to-VLA Adaptation
图 1 · 摘自论文原文
  • 通过对比不同模块参数剪枝影响,定位关键功能模块
  • 在无微调条件下实现12%~30%参数压缩,保留90%原始性能
  • 适用于资源受限环境下高效机器人策略部署

视觉-语言-动作(VLA)模型通过融合预训练视觉-语言模型(VLM)的强大表征,在具身智能领域取得显著进展。然而,其庞大的参数量带来沉重计算负担,且对参数剪枝极为敏感。当前范式常将性能下降视为必然,依赖微调或低秩修正恢复效能。本文挑战这一观点,质疑被剪除参数是否真为冗余——若剪枝后仍需恢复性能才有效,则可能掩盖了对关键参数的盲目删除。通过分析从VLM到VLA适配过程中的参数偏差空间分布,揭示模块间结构化差异。进而引入受控剪枝作为诊断工具:在不进行任何微调的前提下,比较不同参数子集移除对性能的影响,建立适配引发的偏差信号与功能贡献之间的因果关系。基于发现的模块异质性,设计多模块联合剪枝方案。在LIBERO基准测试中,该方法使OpenVLA和$π_{0.5}$模型参数减少12%–30%,同时保持约90%原始性能,而现有剪枝标准在此无恢复条件下导致性能完全崩溃。研究揭示了VLA适配中的参数演化机制,为资源受限环境下的高效、鲁棒机器人策略部署提供新路径。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have made significant strides in embodied intelligence by integrating the powerful representations of pre-trained Vision-Language Models (VLMs). However, the massive parameter scale of VLAs imposes a heavy computational burden, and these models exhibit extreme sensitivity to parameter pruning. Current paradigms often treat the resulting performance degradation as inevitable, relying on fine-tuning or low-rank corrections to recover efficacy. We challenge this convention by questioning whether the removed parameters are truly redundant if VLA pruning necessitates performance recovery to be effective, or if this paradigm masks the indiscriminate pruning of critical parameters. We revisit parameter redundancy through the lens of VLM-to-VLA adaptation, first quantifying the spatial distribution of parameter divergence during adaptation to reveal structured patterns across different modules. Subsequently, we introduce controlled pruning as a diagnostic probe: by comparing the direct impact of removing different parameter subsets on VLA performance without any fine-tuning, we establish a causal link between adaptation-induced divergence signals and functional contributions. Based on the discovered modular heterogeneities, we design a multi-module joint pruning scheme. Evaluations on the LIBERO benchmark demonstrate that our approach reduces the parameters of OpenVLA and $π_{0.5}$ by 12\%--30\% while maintaining approximately 90\% of the original performance without any post-pruning recovery. In contrast, existing parameter pruning criteria result in total performance collapse when evaluated under the same recovery-free constraints. Our study reveals the parameter evolution mechanism in VLA adaptation and provides a new path for deploying efficient, robust robotic policies in resource-constrained environments.

模型剪枝机器人多模态效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。