提出新型二值化视觉变换器,提升边缘设备部署性能。
High-Fidelity Differential-information Driven Binary Vision Transformer
- 用差分信息增强注意力模块,减少二值化信息损失。
- 通过离散哈尔小波分解保留高低频特征相似性,精度显著提升。
- 适合对模型压缩与推理效率要求高的边缘视觉应用。
视觉变换器(ViTs)的二值化为解决高计算/存储需求与边缘设备部署限制之间的矛盾提供了有前景的途径。然而,现有二值化方法常导致性能严重下降或依赖大量全精度模块。为此,我们提出DIDB-ViT,一种高保真且保持原始ViT架构与计算效率的新型二值化视觉变换器。具体而言,设计了融合差分信息的感知注意力模块,以缓解二值化带来的信息丢失并增强高频特征保留;为保持二值化查询(Q)与键(K)张量间相似性的保真度,采用离散哈尔小波进行频率分解,并整合多频段相似性;此外,引入改进的RPReLU激活函数以重构激活分布,扩展模型表达能力。实验表明,DIDB-ViT在多个ViT架构上显著优于现有先进量化方法,在图像分类与分割任务中均取得更优性能。
原文摘要 · Abstract (English)
The binarization of vision transformers (ViTs) offers a promising approach to addressing the trade-off between high computational/storage demands and the constraints of edge-device deployment. However, existing binary ViT methods often suffer from severe performance degradation or rely heavily on full-precision modules. To address these issues, we propose DIDB-ViT, a novel binary ViT that is highly informative while maintaining the original ViT architecture and computational efficiency. Specifically, we design an informative attention module incorporating differential information to mitigate information loss caused by binarization and enhance high-frequency retention. To preserve the fidelity of the similarity calculations between binary Q and K tensors, we apply frequency decomposition using the discrete Haar wavelet and integrate similarities across different frequencies. Additionally, we introduce an improved RPReLU activation function to restructure the activation distribution, expanding the model's representational capacity. Experimental results demonstrate that our DIDB-ViT significantly outperforms state-of-the-art network quantization methods in multiple ViT architectures, achieving superior image classification and segmentation performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。