提出视觉语言模型层跳过理论条件,实现高效推理而不降性能。
Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
- 基于可验证的冗余性指标,建立层跳过的统一理论框架。
- 实证发现早期和晚期视觉模态信息均存在冗余,可安全跳过。
- 为模型压缩提供理论依据,适合关注高效推理的研究者。
视觉语言模型在多种任务中表现卓越,但其庞大的规模导致推理成本高昂。近期研究发现多模态处理中存在显著冗余,可通过跳过某些层实现高效推理而性能损失极小。然而现有剪枝技术仍依赖启发式方法或超参数调优,缺乏可解释的理论判据。本文提出一个统一框架,刻画了在何种冗余条件下剪枝可提升效率且不损害性能。核心是实验可验证、可解释的冗余度量,无需依赖下游任务表现即可评估。应用该框架,我们复现了先前结论:不同模型中早期和晚期视觉令牌均具冗余性,并验证了理论条件与实际性能下降的一致性。此外,本框架提供了对视觉语言模型冗余性的理论理解,整合了现代层跳过技术的核心思想。
原文摘要 · Abstract (English)
Vision-language models achieve incredible performance across a wide range of tasks, but their large size makes inference costly. Recent work has shown that multimodal processing contains significant redundancies, making it possible to skip certain layers with minimal performance loss. Yet current pruning techniques remain ad-hoc, relying on heuristics or hyperparameter sweeps rather than principled criteria for determining when layer skipping is beneficial. In this paper, we propose a unified framework that characterizes the redundancy conditions under which pruning can enhance efficiency without sacrificing performance. Central to our approach are experimentally verifiable and interpretable notions of redundancy that can be evaluated without requiring downstream task performance as a metric. Applying this framework, we corroborate prior findings that both early and late vision tokens are redundant across models, and we validate our conditions by showing they align with actual performance degradation. Beyond these empirical results, our framework provides a theoretically grounded understanding of redundancy in VLMs and unifies many of the ideas behind modern layer-skipping techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。