让视觉模型关注被忽略的物体,提升图文匹配准确性
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
- 用注意力引导模块增强视觉编码器对非显著物体的感知
- 在多个数据集上实现最高检索准确率,显著提升小物体识别能力
- 适合需要精准视觉理解的多模态任务,如图像描述生成
多模态大语言模型在文本到视频生成、视觉问答等应用中取得进展,依赖视觉编码器将非文本数据转为向量表示。然而现有编码器或缺乏语义对齐,或忽略非显著物体。本文提出GiVE方法,通过注意力引导适配器(AG-Adapter)和面向物体的视觉语义学习模块,引入三种新损失函数:面向物体的图文对比(OITC)、物间图像对比(OIIC)和物体判别(OID),有效提升物体关注度、检索准确性和表征全面性。贡献包括动态视觉焦点调节、新型损失函数设计及多物体指令数据集(MOInst)。实验表明该方法达到当前最优性能。
原文摘要 · Abstract (English)
Multimodal Large Language Models have advanced AI in applications like text-to-video generation and visual question answering. These models rely on visual encoders to convert non-text data into vectors, but current encoders either lack semantic alignment or overlook non-salient objects. We propose the Guiding Visual Encoder to Perceive Overlooked Information (GiVE) approach. GiVE enhances visual representation with an Attention-Guided Adapter (AG-Adapter) module and an Object-focused Visual Semantic Learning module. These incorporate three novel loss terms: Object-focused Image-Text Contrast (OITC) loss, Object-focused Image-Image Contrast (OIIC) loss, and Object-focused Image Discrimination (OID) loss, improving object consideration, retrieval accuracy, and comprehensiveness. Our contributions include dynamic visual focus adjustment, novel loss functions to enhance object retrieval, and the Multi-Object Instruction (MOInst) dataset. Experiments show our approach achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。