arXiv:2410.13666cs.CVcs.CL2024-10被引 1

构建首个综合多模态推理任务的基准,挑战AI跨视觉与语言理解能力。

VL-GLUE: A Suite of Fundamental yet Challenging Visuo-Linguistic Reasoning Tasks

  • 设计7个需视觉与文本联合推理的任务,覆盖多样图像与领域文本。
  • 包含超10万样本,涵盖合成图、日常场景、图表及课程等真实数据。
  • 现有大模型在此基准表现不佳,推动更鲁棒多模态系统发展。

从异构输入(如图像、文本、音频)中进行推断是人类完成日常任务的重要能力,也是先进人工智能系统所需的关键能力。尽管当前顶尖模型在单一计算机视觉或自然语言处理任务上已接近人类水平,但在需要跨视觉与文本联合推理的任务上仍表现不足。受GLUE(Wang et al., 2018)——自然语言理解多任务基准的启发,本文提出VL-GLUE。该基准包含超过10万个样本,覆盖7个核心需多模态推理的任务。数据涵盖合成图像、日常生活场景、图表与复杂示意图,以及烹饪、政治、体育和中学课程等多样化领域文本,体现现实世界对多模态理解的需求。实验表明,现有大规模视觉-语言模型在此基准上表现有限,亟需发展具备强健多模态推理能力的新系统。

原文摘要 · Abstract (English)

Deriving inference from heterogeneous inputs (such as images, text, and audio) is an important skill for humans to perform day-to-day tasks. A similar ability is desirable for the development of advanced Artificial Intelligence (AI) systems. While state-of-the-art models are rapidly closing the gap with human-level performance on diverse computer vision and NLP tasks separately, they struggle to solve tasks that require joint reasoning over visual and textual modalities. Inspired by GLUE (Wang et. al., 2018)- a multitask benchmark for natural language understanding, we propose VL-GLUE in this paper. VL-GLUE consists of over 100k samples spanned across seven different tasks, which at their core require visuo-linguistic reasoning. Moreover, our benchmark comprises of diverse image types (from synthetically rendered figures, and day-to-day scenes to charts and complex diagrams) and includes a broad variety of domain-specific text (from cooking, politics, and sports to high-school curricula), demonstrating the need for multi-modal understanding in the real-world. We show that this benchmark is quite challenging for existing large-scale vision-language models and encourage development of systems that possess robust visuo-linguistic reasoning capabilities.

多模态推理任务视觉语言基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。