让多模态大模型更精准理解指令中的视觉细节
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
- 在视觉编码早期融合文本指令,提升相关特征提取
- 通过过滤冗余信息,训练成本降低显著
- 适配任意解码器架构,适合视觉任务主导场景
多模态大语言模型虽快速逼近人类视觉感知能力,但在关注细微图像特征或精确定位小物体等方面仍显不足。现有方案多依赖多视觉编码器或高分辨率原图处理,而极少研究将文本指令用于增强视觉表征,导致在视觉主导任务中注意力分散,我们称此现象为‘弱视’(Amblyopia)。本文提出 Panther,一种紧密遵循用户指令、能精准定位目标的多模态大模型,如黑豹般敏捷。其由三部分构成:Panther-VE 在视觉编码早期融合指令信息,提取最相关视觉表征;Panther-Bridge 具备强大过滤能力,大幅减少冗余信息,显著降低训练成本;Panther-Decoder 可与任意解码器仅架构兼容。实验结果表明,该模型在视觉主导基准上表现优异。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) are closing the gap to human visual perception capability rapidly, while, still lag behind on attending to subtle images details or locating small objects precisely, etc. Common schemes to tackle these issues include deploying multiple vision encoders or operating on original high-resolution images. Few studies have concentrated on taking the textual instruction into improving visual representation, resulting in losing focus in some vision-centric tasks, a phenomenon we herein termed as Amblyopia. In this work, we introduce Panther, a MLLM that closely adheres to user instruction and locates targets of interests precisely, with the finesse of a black panther. Specifically, Panther comprises three integral components: Panther-VE, Panther-Bridge, and Panther-Decoder. Panther-VE integrates user instruction information at the early stages of the vision encoder, thereby extracting the most relevant and useful visual representations. The Panther-Bridge module, equipped with powerful filtering capabilities, significantly reduces redundant visual information, leading to a substantial savings in training costs. The Panther-Decoder is versatile and can be employed with any decoder-only architecture of LLMs without discrimination. Experimental results, particularly on vision-centric benchmarks, have demonstrated the effectiveness of Panther.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。