arXiv:2412.10151cs.CVcs.AI2024-12被引 3

构建多语言视觉问答基准,测试模型从多个文本中选有用信息的能力

VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation

  • 设计五段落输入的VQA数据集,测试模型筛选有效信息的能力
  • 生成3.2万条指令跟随样本,提升视觉语言模型的检索增强生成能力
  • 适合作为评估大模型多源知识整合能力的基准,尤其适合研究RAG的学者

我们提出了VLR-Bench,一个基于检索增强生成(RAG)的视觉语言模型(VLM)评估基准。与以往依赖外部知识的视觉问答数据集不同,VLR-Bench包含五个输入段落,能够测试模型判断哪个段落对回答问题有用的能力,这是此前研究中缺失的关键能力。为此,我们构建了一个包含32,000条自动生成的指令跟随样本的数据集,称为VLR-IF。该数据集旨在通过让模型学习根据输入段落生成恰当答案,从而增强其RAG能力。我们使用最先进的基于Llama3的VLM——Llava-Llama-3模型,验证了该基准和训练数据的有效性。VLR-Bench与VLR-IF数据集已公开发布。

原文摘要 · Abstract (English)

We propose the VLR-Bench, a visual question answering (VQA) benchmark for evaluating vision language models (VLMs) based on retrieval augmented generation (RAG). Unlike existing evaluation datasets for external knowledge-based VQA, the proposed VLR-Bench includes five input passages. This allows testing of the ability to determine which passage is useful for answering a given query, a capability lacking in previous research. In this context, we constructed a dataset of 32,000 automatically generated instruction-following examples, which we denote as VLR-IF. This dataset is specifically designed to enhance the RAG capabilities of VLMs by enabling them to learn how to generate appropriate answers based on input passages. We evaluated the validity of the proposed benchmark and training data and verified its performance using the state-of-the-art Llama3-based VLM, the Llava-Llama-3 model. The proposed VLR-Bench and VLR-IF datasets are publicly available online.

视觉问答检索生成多语言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。