发现视觉令牌剪枝在复杂推理中失效原因并提出改进方案
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding

- 识别出解码阶段视觉信息需求动态变化是剪枝失败主因
- 新方法DSTP无需训练,显著提升复杂推理任务表现
- 适配多种主流模型,计算开销极低且效果稳定
近期视觉令牌剪枝被用于处理多模态大语言模型中的海量视觉令牌。然而我们发现,现有剪枝方法在简单视觉理解任务中表现可靠,却难以泛化到复杂视觉推理任务,这一关键差距此前未被充分研究。通过系统分析,我们识别出解码过程中相关视觉信息的动态偏移(RVIS)是导致失败的主要原因。为此,我们提出解码阶段感知的剪枝框架DSTP,一种无需训练的附加机制,使现有剪枝方法能对齐解码阶段不断变化的推理需求。大量实验表明,DSTP显著缓解了剪枝方法在复杂推理任务中的性能下降,同时在各类视觉理解基准上持续带来性能提升。此外,DSTP在多种先进架构上均有效,展现出良好的通用性与极低的计算开销。
原文摘要 · Abstract (English)
Recently, visual token pruning has been studied to handle the vast number of visual tokens in Multimodal Large Language Models. However, we observe that while existing pruning methods perform reliably on simple visual understanding, they struggle to effectively generalize to complex visual reasoning tasks, a critical gap underexplored in previous studies. Through a systematic analysis, we identify Relevant Visual Information Shift (RVIS) during decoding as the primary failure driver. To address this, we propose Decoding-stage Shift-aware Token Pruning (DSTP), a training-free add-on framework that enables existing pruning methods to align visual tokens with shifting reasoning requirements during the decoding stage. Extensive experiments demonstrate that DSTP significantly mitigates performance degradation of pruning methods in complex reasoning tasks, while consistently yielding performance gains even across visual understanding benchmarks. Furthermore, DSTP demonstrates effectiveness across diverse state-of-the-art architectures, highlighting its generalizability and efficiency with minimal computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。