用视觉特征引导的渐进式令牌剪枝,显著降低模型计算量。
Back to Fundamentals: Low-Level Visual Features Guided Progressive Token Pruning
- 基于多尺度熵和低层视觉特征设计动态评分机制
- 在多个数据集上实现20%-45%计算量减少,边缘区域精度稳定
- 无需修改结构或训练,可直接集成到现有模型中
视觉变换器(ViTs)在语义分割中表现优异,但计算开销大,难以部署于资源受限设备。现有令牌剪枝方法常忽略基础视觉数据特性。本文提出LVTP框架,通过多尺度Tsallis熵与低层视觉特征双聚类,融合高层语义与底层视觉属性以实现精准分割。创新性地采用多尺度Tsallis熵加权的动态评分机制,克服传统单参数熵的局限。同时引入低层特征分析,在优化计算成本的同时保留关键边缘信息。作为即插即用模块,无需架构修改或额外训练。在多个数据集上的评估显示,计算量降低20%-45%,性能损失可忽略,尤其在复杂边缘区域优于现有方法。
原文摘要 · Abstract (English)
Vision Transformers (ViTs) excel in semantic segmentation but demand significant computation, posing challenges for deployment on resource-constrained devices. Existing token pruning methods often overlook fundamental visual data characteristics. This study introduces 'LVTP', a progressive token pruning framework guided by multi-scale Tsallis entropy and low-level visual features with twice clustering. It integrates high-level semantics and basic visual attributes for precise segmentation. A novel dynamic scoring mechanism using multi-scale Tsallis entropy weighting overcomes limitations of traditional single-parameter entropy. The framework also incorporates low-level feature analysis to preserve critical edge information while optimizing computational cost. As a plug-and-play module, it requires no architectural changes or additional training. Evaluations across multiple datasets show 20%-45% computational reductions with negligible performance loss, outperforming existing methods in balancing cost and accuracy, especially in complex edge regions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。