首篇系统综述多模态大模型的视觉提示方法。
Visual Prompting in Multimodal Large Language Models: A Survey
- 梳理视觉提示、生成与学习的分类体系
- 涵盖图像提示自动生成与跨模态对齐技术
- 适合研究多模态推理与提示工程的学者
多模态大语言模型(MLLMs)为预训练大语言模型(LLMs)赋予了视觉理解能力。尽管文本提示在LLMs中已得到广泛研究,视觉提示因其更细粒度和自由形式的指令表达而逐渐兴起。本文首次全面综述了MLLM中的视觉提示方法,聚焦于视觉提示、提示生成、组合推理与提示学习。我们对现有视觉提示进行分类,并讨论图像上自动提示标注的生成方法。同时,探讨提升视觉编码器与主干LLM之间对齐的视觉提示技术,涉及视觉定位、物体指代与组合推理能力。此外,总结了模型训练与上下文学习方法,以增强MLLM对视觉提示的感知与理解。本文系统分析了当前视觉提示方法,并展望其未来发展。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) equip pre-trained large-language models (LLMs) with visual capabilities. While textual prompting in LLMs has been widely studied, visual prompting has emerged for more fine-grained and free-form visual instructions. This paper presents the first comprehensive survey on visual prompting methods in MLLMs, focusing on visual prompting, prompt generation, compositional reasoning, and prompt learning. We categorize existing visual prompts and discuss generative methods for automatic prompt annotations on the images. We also examine visual prompting methods that enable better alignment between visual encoders and backbone LLMs, concerning MLLM's visual grounding, object referring, and compositional reasoning abilities. In addition, we provide a summary of model training and in-context learning methods to improve MLLM's perception and understanding of visual prompts. This paper examines visual prompting methods developed in MLLMs and provides a vision of the future of these methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。