arXiv:2501.03675cs.CV2025-01被引 3

用合成数据提升多图推理能力,成本低效果好

SMIR: Efficient Synthetic Data Pipeline To Improve Multi-Image Reasoning

  • 通过多模态嵌入提取关联图像,结合大模型生成高质量指令
  • 构建16万条合成数据,支持复杂多图任务训练
  • 提供可评估的多轮推理基准,适合研究多模态推理的学者

视觉语言模型(VLMs)在单图理解上表现优异,得益于高质量指令数据集。然而,由于两个关键挑战:(1) 构建包含相关图像和复杂推理指令的数据集资源消耗大;(2) 缺乏可靠的多图任务评估基准,多图推理在开源社区仍处于探索阶段。为此,我们提出SMiR——一种高效的多图推理合成数据生成流水线,并基于此生成高质量数据集。SMiR通过多模态嵌入高效提取相关图像,融合视觉与描述信息,并利用开源大语言模型生成优质指令。该方法生成了16万条合成训练样本,为闭源方案提供了低成本替代方案。此外,我们提出SMiR-Bench,一个包含200个多样化样本、涵盖七类复杂推理任务的多图推理评估基准。该基准支持多轮交互,采用VLM裁判评估自由文本输出,全面评估模型跨模态表达力与推理能力。通过微调开源VLM并在SMiR-Bench上评估,验证了SMiR的有效性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) excel at understanding single images, aided by high-quality instruction datasets. However, multi-image reasoning remains underexplored in the open-source community due to two key challenges: (1) scaling datasets with correlated images and complex reasoning instructions is resource-intensive, and (2) robust evaluation benchmarks for multi-image tasks are lacking. To address this, we introduce SMiR, a synthetic data-generation pipeline for multi-image reasoning, along with a high-quality dataset generated using this pipeline. SMiR efficiently extracts correlated images via multimodal embeddings, integrates visual and descriptive information, and leverages open-source LLMs to generate quality instructions. Using this approach, we produce 160K synthetic training samples, offering a cost-effective alternative to closed-source solutions. Additionally, we present SMiR-Bench, a multi-image reasoning benchmark comprising 200 diverse examples across seven complex reasoning tasks. SMiR-Bench is multi-turn and employs a VLM judge to evaluate free-form responses, providing a comprehensive assessment of model expressiveness and reasoning capability across modalities. We demonstrate the effectiveness of SMiR by fine-tuning open-source VLMs and evaluating them on SMiR-Bench.

多图推理合成数据视觉语言模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。