让视觉语言模型通过联网搜索,识别从未见过的图像内容。
Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines
- 用视觉语言模型配合网络代理,实现跨网页检索增强生成。
- 在开放集和封闭集评测中均显著超越现有模型。
- 适合需要实时理解新图像的场景,如跨领域问答。
搜索引擎能通过文本检索未知信息。然而,传统方法在理解陌生视觉内容时表现不佳,例如识别模型从未见过的物体。这一挑战在大型视觉语言模型(VLMs)中尤为突出:若图像中的物体未在训练数据中出现,模型难以生成可靠回答。此外,随着新物体和事件不断涌现,频繁更新VLMs因计算开销过大而不可行。为此,我们提出视觉搜索助手(Vision Search Assistant),一种新型框架,促进VLMs与网络代理协作。该方法结合VLMs的视觉理解能力与网络代理的实时信息获取能力,通过网络实现开放世界的检索增强生成。通过整合视觉与文本表征,模型即使面对系统未曾见过的图像也能提供有依据的回答。在开放集和封闭集问答基准上的大量实验表明,该方法显著优于其他模型,并可广泛适配现有VLMs。
原文摘要 · Abstract (English)
Search engines enable the retrieval of unknown information with texts. However, traditional methods fall short when it comes to understanding unfamiliar visual content, such as identifying an object that the model has never seen before. This challenge is particularly pronounced for large vision-language models (VLMs): if the model has not been exposed to the object depicted in an image, it struggles to generate reliable answers to the user's question regarding that image. Moreover, as new objects and events continuously emerge, frequently updating VLMs is impractical due to heavy computational burdens. To address this limitation, we propose Vision Search Assistant, a novel framework that facilitates collaboration between VLMs and web agents. This approach leverages VLMs' visual understanding capabilities and web agents' real-time information access to perform open-world Retrieval-Augmented Generation via the web. By integrating visual and textual representations through this collaboration, the model can provide informed responses even when the image is novel to the system. Extensive experiments conducted on both open-set and closed-set QA benchmarks demonstrate that the Vision Search Assistant significantly outperforms the other models and can be widely applied to existing VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。