针对多模态视觉模型中2D与3D信息重要性差异,提出分阶段剪枝框架提升效率。
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness

- 按多模态数据处理流程设计三阶段分析,捕捉2D/3D信息重要性动态变化
- 实现最高2.55倍推理加速,精度损失极小,额外开销仅5.8%
- 专为2D+3D融合的多视觉模态模型优化,适合追求高效部署的研究者
视觉-语言-动作(VLA)模型已成为具身智能主流。近期模型从纯2D输入扩展到2D+3D多模态输入,形成多视觉模态VLA(MVLA)模型。尽管提升了空间感知能力,但模态扩展导致输入标记数量激增,带来更大加速需求。标记剪枝是针对MVLA模型的有效优化方法。然而,现有剪枝方案仍针对纯2D VLA模型设计,忽略2D与3D模态在信息重要性上的差异。本文遵循多模态数据在MVLA模型中的应用流程,提出三阶段分析以捕捉2D/3D模态重要性的差异与动态变化。基于此,构建相应的三阶段标记剪枝框架,实现最优2D/3D标记选择与高效剪枝。实验表明,该框架在仅增加5.8%开销的前提下,实现最高2.55倍的推理加速,且精度损失极小。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have emerged as the mainstream of embodied intelligence. Recent VLA models have expanded their input modalities from 2D-only to 2D+3D paradigms, forming multi-visual-modal VLA (MVLA) models. Despite achieving improved spatial perception, MVLA faces a greater acceleration demand due to the increased number of input tokens caused by modal expansion. Token pruning is an effective optimization methods tailored to MVLA models. However, existing token pruning schemes are designed for 2D-only VLA models, ignoring 2D/3D modality salience differences. In this paper, we follow the application process of multi-modal data in MVLA models and develop a tri-stage analysis to capture the discrepancy and dynamics of 2D/3D modality salience. Based on these, we propose a corresponding tri-stage token pruning framework for MVLA models to achieve optimal 2D/3D token selection and efficient pruning. Experiments show that our framework achieves up to a 2.55x inference speedup with minimal accuracy loss, while only costing 5.8% overhead. Our Code is coming soon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。