通过感知领域差异选择剪枝层,提升视觉语言模型稳定性。
Understanding Pruning Regimes in Vision-Language Models Through Domain-Aware Layer Selection
- 基于领域激活相似性识别关键层,实现分域剪枝
- 低剪枝率下性能对剪枝层敏感,高剪枝率下结构连续性主导
- 适用于需数学推理的视觉语言模型轻量化设计
基于Transformer的视觉语言模型(VLMs)存在显著深度冗余,但特定解码器层的移除影响尚不明确,尤其在感知与多步推理强耦合的领域。本文通过领域感知激活相似性分析,衡量各层在数学与非数学输入下的表征变换强度,提出数学感知、非数学感知及混合排序标准,识别目标领域内输入输出激活变化最小的层。在两个先进VLMs和涵盖数学与通用多模态任务的广泛基准上,发现一致的三阶段剪枝规律:低剪枝预算下性能高度依赖被剪层;中等预算时方法趋于收敛,因结构损伤累积;高预算时结构连续性占优,青睐间距感知策略。领域感知排序在排名敏感阶段表现最优,且在大剪枝量下匹配或超越结构感知基线。结果揭示了深度对领域特异性行为的影响,并提供一种可解释、实用的模型降深方法,兼顾数学与通用能力。
原文摘要 · Abstract (English)
Transformer-based vision-language models (VLMs) contain substantial depth redundancy, yet the effect of removing specific decoder layers remains poorly understood, especially for domains that require tight coupling between perception and multi-step reasoning. We study structured decoder layer pruning through the lens of domain-aware activation similarity, measuring how strongly each layer transforms representations for math versus non-math inputs. This yields simple math-aware, non-math-aware, and mixed ranking criteria that identify layers whose input-output activations change least within a target domain. Across two state-of-the-art VLMs and a broad suite of math and general multimodal benchmarks, we uncover a consistent three-regime structure: at low pruning budgets, performance is highly sensitive to which layers are removed; at moderate budgets, methods converge as structural damage accumulates; and at high budgets, structural continuity dominates, favoring spacing-aware strategies. Our domain-aware rankings achieve the strongest stability in the ranking-sensitive regime, while matching or exceeding structure-aware baselines at larger budgets. These results provide a clearer picture of how depth contributes to domain-specific behavior in VLMs and offer a practical, interpretable approach to reducing model depth without sacrificing essential mathematical or general vision-language capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。