arXiv:2503.10582cs.CVcs.AI2025-03EMNLP被引 39

用网页搜索构建90万条多模态推理数据,提升视觉语言模型理解力

VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search

  • 通过谷歌图片搜索拓展种子图像,爬取超70万网页构建数据集
  • 40%为图文问答对,微调后在多个基准上提升10-20个百分点
  • 适合做多模态推理、跨学科知识理解的模型研究者使用

视觉语言模型在感知任务上已取得显著进展,但在推理任务上的表现受限于高质量、多样化的训练数据匮乏。本文提出VisualWebInstruct,一种利用搜索引擎构建多样化高质多模态数据集的新方法,覆盖数学、物理、金融、化学等多个领域。基于3万张精选种子图像,通过Google Image Search定位含相似图像的网站,从超过70万唯一URL中收集并处理HTML数据。经内容提取、过滤与合成,构建约90万条问答对,其中40%为图文问答对,其余为纯文本问答对。在VisualWebInstruct上微调的模型性能显著提升:在Llava-OV上各基准提升10-20个绝对点,在MAmmoTH-VL基础上微调获得5个绝对点增益。最佳模型MAmmoTH-VL2在10B参数级别达到当前最优,在MMMU-Pro(40.7)、MathVerse(42.6)和DynaMath(55.7)上表现突出。结果验证了该数据集在增强视觉语言模型复杂多模态推理能力方面的有效性。

原文摘要 · Abstract (English)

Vision-Language Models have made significant progress on many perception-focused tasks. However, their progress on reasoning-focused tasks remains limited due to the lack of high-quality and diverse training data. In this work, we aim to address the scarcity of reasoning-focused multimodal datasets. We propose VisualWebInstruct, a novel approach that leverages search engines to create a diverse and high-quality dataset spanning multiple disciplines, including mathematics, physics, finance, and chemistry, etc. Starting with a meticulously selected set of 30,000 seed images, we employ Google Image Search to identify websites containing similar images. We collect and process HTML data from over 700K unique URLs. Through a pipeline of content extraction, filtering, and synthesis, we construct a dataset of approximately 900K question-answer (QA) pairs, with 40% consisting of visual QA pairs and the remaining comprising text-based QA pairs. Models fine-tuned on VisualWebInstruct demonstrate significant performance improvements: (1) fine-tuning on Llava-OV results in 10-20 absolute points improvement across benchmarks, and (2) fine-tuning from MAmmoTH-VL yields a 5 absolute points gain across benchmarks. Our best model, MAmmoTH-VL2, achieves state-of-the-art performance within the 10B parameter class on MMMU-Pro (40.7), MathVerse (42.6), and DynaMath (55.7). These results highlight the effectiveness of our dataset in enhancing the reasoning capabilities of vision-language models for complex multimodal tasks.

多模态推理数据构建视觉语言模型网页搜索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。