构建大规模合成视觉逻辑题数据集,提升模型推理能力
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
- 基于规则生成图像,确保题目与答案严格对应
- 在逻辑推理任务上性能显著提升,跨任务泛化效果好
- 适合需要强逻辑推理的多模态研究者使用
视觉语言模型(VLMs)需具备有效的多模态推理和逻辑一致决策能力,这对图表理解与空间问题求解至关重要。然而当前VLM推理缺乏大规模且结构清晰的训练数据。为此,我们提出VisualSphinx,首个大规模合成视觉逻辑推理训练数据集。为解决带精确答案的图像合成难题,我们设计了从种子问题提取并扩展谜题规则的规则到图像合成流程,并生成用于拼装谜题样本的合成代码。实验表明,在VisualSphinx上使用GRPO训练的VLM受益于数据集的逻辑连贯性和可读性,在逻辑推理任务中表现更优。由此获得的推理能力还提升了代数、算术与几何推理等其他任务的表现。
原文摘要 · Abstract (English)
Vision language models (VLMs) are expected to perform effective multimodal reasoning and make logically coherent decisions, which is critical to tasks such as diagram understanding and spatial problem solving. However, current VLM reasoning lacks large-scale and well-structured training datasets. To bridge this gap, we propose VisualSphinx, a first-of-its-kind large-scale synthetic visual logical reasoning training data. To tackle the challenge of image synthesis with grounding answers, we propose a rule-to-image synthesis pipeline, which extracts and expands puzzle rules from seed questions and generates the code of grounding synthesis image synthesis for puzzle sample assembly. Experiments demonstrate that VLM trained using GRPO on VisualSphinx benefit from logical coherence and readability of our dataset and exhibit improved performance on logical reasoning tasks. The enhanced reasoning capabilities developed from VisualSphinx also benefit other reasoning tasks such as algebraic reasoning, arithmetic reasoning and geometry reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。