用视觉注意力机制精简图像令牌,提升多模态大模型效率
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
- 模仿人类看图答题习惯,按CLIP度量筛选关键图像令牌
- 在12个数据集上实现计算开销显著降低,性能保持稳定
- 适合追求高效部署的多模态模型开发者
多模态大语言模型(MLLMs)在多个领域表现卓越,但其资源消耗急剧上升。本文提出一种名为基于CLIP度量的令牌压缩方法(TRIM),旨在提升MLLMs效率而不损失性能。受视觉问答任务中人类注意力模式启发,TRIM从图像令牌中智能筛选关键信息。该方法在12个数据集上进行了广泛测试,结果表明在大幅降低计算开销的同时,性能保持一致水平。这项研究为高效多模态大模型的发展迈出关键一步,推动高性能模型的可及性与可持续性。
原文摘要 · Abstract (English)
The rapid advancement of Multimodal Large Language Models (MLLMs) has led to remarkable performances across various domains. However, this progress is accompanied by a substantial surge in the resource consumption of these models. We address this pressing issue by introducing a new approach, Token Reduction using CLIP Metric (TRIM), aimed at improving the efficiency of MLLMs without sacrificing their performance. Inspired by human attention patterns in Visual Question Answering (VQA) tasks, TRIM presents a fresh perspective on the selection and reduction of image tokens. The TRIM method has been extensively tested across 12 datasets, and the results demonstrate a significant reduction in computational overhead while maintaining a consistent level of performance. This research marks a critical stride in efficient MLLM development, promoting greater accessibility and sustainability of high-performing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。