用合成数据和后训练让小模型搞定抽象视觉推理,超越大模型。
On Data Synthesis and Post-training for Visual Abstract Reasoning
- 自动生成抽象视觉推理数据,分步训练模型
- 7B模型在基准上大幅超越Qwen/GPT-4o等大模型
- 不损失通用多模态理解能力,适合研究抽象推理
本文首次尝试解决大视觉语言模型(VLMs)在抽象视觉推理(AVR)任务中的难题。我们使一个通用的LLaVA-NeXT 7B模型具备感知与推理特定AVR问题的能力,在代表性基准上显著超越开源(如Qwen-2-VL-72B)和闭源强大模型(如GPT-4o)。此前几乎所有VLM在该任务上表现失败或接近随机。关键成功在于创新的数据合成与后训练流程,逐步降低任务难度,引导模型有效学习。该7B模型在保持通用多模态理解能力的同时,展现出优秀的AVR性能。本工作为该领域提供早期范式,有望激发更多研究。
原文摘要 · Abstract (English)
This paper is a pioneering work attempting to address abstract visual reasoning (AVR) problems for large vision-language models (VLMs). We make a common LLaVA-NeXT 7B model capable of perceiving and reasoning about specific AVR problems, surpassing both open-sourced (e.g., Qwen-2-VL-72B) and closed-sourced powerful VLMs (e.g., GPT-4o) with significant margin. This is a great breakthrough since almost all previous VLMs fail or show nearly random performance on representative AVR benchmarks. Our key success is our innovative data synthesis and post-training process, aiming to fully relieve the task difficulty and elicit the model to learn, step by step. Our 7B model is also shown to be behave well on AVR without sacrificing common multimodal comprehension abilities. We hope our paper could serve as an early effort in this area and would inspire further research in abstract visual reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。