arXiv:2605.17499cs.LG2026-05中稿 · ICASSP 2026

用文本引导提前退出机制,降低CLIP图像编码器的计算开销

t-gems: text-guided exit modules for decreasing clip image encoder

论文配图:t-gems: text-guided exit modules for decreasing clip image encoder
图 1 · 摘自论文原文
  • 基于文本描述分析中间层语义,设计文本引导退出模块
  • 在保持跨模态理解性能前提下,有效减少编码器计算成本
  • 适合需要高效多模态推理的应用场景

多模态深度神经网络通过融合不同模态数据提升深层理解能力。不同模态数据通常被映射到共享潜在空间以进行相似性计算,但这一过程因大型图像编码器及测试数据的全程处理而资源消耗巨大。早期退出方法通过利用中间层特征降低计算负载,节省时间和内存。然而,针对图像-文本等多模态数据设计此类方法仍具挑战性。本研究分析了如CLIP等编码器中间层中的语义内容分布,并发现其可由文本描述推导。为此,提出文本引导退出模块(T-GEMs)与基于速率的正则化器,实现对编码器使用成本的可控调节,同时维持跨模态理解性能。

原文摘要 · Abstract (English)

Multimodal deep neural networks enhance deep comprehension by integrating diverse data modalities. Data from different modalities are typically projected into a shared latent space for similarity computation, but this process is resource intensive due to large image encoders and equal processing of test data during prediction. Early exit methods reduce computational load by utilizing intermediate layers, saving time and memory. However, developing such methods is challenging for multimodal data like image-text pairs. This study investigates the semantic content distributions present in intermediate layers of encoders such as CLIP, which can be derived from textual descriptions. We introduce Text-Guided Exit Modules (T-GEMs) and a rate-based regularizer to control encoder usage costs while maintaining cross-modal understanding performance.

多模态模型压缩CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。