arXiv:2603.06270cs.CVcs.AI2026-03

提出分层偏好调控剪枝框架,实现视觉语言模型高效压缩与抗幻觉平衡。

HiPP-Prune: Hierarchical Preference-Conditioned Structured Pruning for Vision-Language Models

论文配图:HiPP-Prune: Hierarchical Preference-Conditioned Structured Pruning for Vision-Language Models
图 1 · 摘自论文原文
  • 基于多目标条件资源分配,分层决策剪枝策略与层间预算分配。
  • 在相同稀疏度下,相比基线降低37%物体幻觉,提升任务性能与鲁棒性。
  • 支持用户自定义偏好,适合部署需兼顾精度与可信度的视觉语言系统。

为实现视觉语言模型(VLMs)高效部署,剪枝面临压缩导致任务效用下降与视觉定位失准的问题,常在相同稀疏度下加剧物体幻觉。本文提出HiPP-Prune,一种分层偏好调控的结构化剪枝框架,将剪枝视为多目标下的条件资源分配。该框架进行计划级决策:单次策略调用输出全局剪枝蓝图,通过分解为整体稀疏度预算与层间分配,支持用户指定偏好向量下的可查询权衡。针对VLM特有失效模式,策略状态融合从视觉标记与语言隐藏状态间注意力流提取的视觉敏感信号,避免过度剪裁促进跨模态融合的关键视觉层。采用计划级群体相对策略优化(GRPO)优化剪枝方案,在包含任务效用、幻觉鲁棒性(POPE)、压缩率及类突触流稳定性代理的多目标回报下,减少高稀疏度下的无效探索。在LLaVA上基于POPE与ScienceQA的实验表明,HiPP-Prune发现多种非占优剪枝方案,可在匹配稀疏度预算下提供可控的鲁棒性-性能权衡。

原文摘要 · Abstract (English)

Pruning vision-language models (VLMs) for efficient deployment is challenging because compression can affect not only task utility but also visual grounding, often amplifying object hallucinations even at the same sparsity level. We present HiPP-Prune, a hierarchical preference-conditioned structured pruning framework that treats pruning as conditional resource allocation under multiple objectives. HiPP-Prune makes plan-level decisions: a single policy invocation outputs a global pruning blueprint by factorizing decisions into an overall sparsity budget and a layer-wise allocation, enabling queryable trade-offs via a user-specified preference vector. To account for VLM-specific failure modes, our policy state integrates a visual sensitivity signal derived from attention flow between vision tokens and language hidden states, discouraging over-pruning of vision-critical layers that facilitate cross-modal fusion. We optimize pruning plans with plan-level Group Relative Policy Optimization (GRPO) under a multi-objective return that combines task utility, hallucination robustness (POPE), compression, and a synaptic-flow-inspired stability proxy to reduce unproductive exploration in high-sparsity regimes. Experiments on LLaVA with POPE and ScienceQA demonstrate that HiPP-Prune discovers diverse non-dominated pruning plans and provides controllable robustness--utility trade-offs under matched sparsity budgets.

视觉语言模型结构化剪枝抗幻觉多目标优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。