用合成数据和强化学习提升视觉模型多图分析能力
SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning

- 构建可控扰动的图像对,生成需跨图推理的任务
- 新方法使模型在多图推理上准确率提升36.95%
- 适合研究多模态推理或想提升模型解释力的开发者
视觉语言模型(VLMs)在感知任务中表现强劲,但在涉及多视觉状态分析的推理任务(如多图对比、变化检测、多步视觉推断)中仍受限。现有评测基准很少同时要求明确的视觉比较与分析推理,导致该能力未被充分探索。为此,本文提出SD-MAR框架,通过可控扰动生成配对视觉场景,并构造涵盖语义变化归因与量化比较的推理任务。采用改进的GRPO-lite结合后向折扣分配(BDA)的强化学习方法训练模型,去除KL正则化以增强策略优化,并将更大奖励分配给形成结论的后期推理步骤。在Qwen2.5-VL-7B和InternVL3-8B上的实验表明,于SD-MAR上微调后,域内准确率最高提升36.95%,其中Qwen2.5-VL-7B性能超越GPT-4.1。域外泛化能力保持稳定:在MME、MMMU-Pro、MathVista上性能波动不超过1%,在MMBench上提升最高达4%。大模型评分显示,两模型在逻辑连贯性与解释质量上均有持续提升。
原文摘要 · Abstract (English)
Vision Language Models (VLMs) demonstrate strong perceptual abilities but remain limited in tasks requiring analytical reasoning across multiple visual states, such as multi-image comparison, change detection, and multi-step visual inference. These capabilities are critical for real-world multimodal applications where reasoning must be grounded in systematic differences between visual contexts. However, existing benchmarks rarely require both explicit visual comparison and analytical reasoning, leaving this capability underexplored. To address this gap, we introduce SD-MAR (Synthetic Data for Multi-image Analytical Reasoning), a framework for training and evaluating VLMs on multi-image analytical reasoning. SD-MAR constructs paired visual scenarios through controlled perturbations and generates reasoning tasks spanning semantic change attribution and quantitative comparison. We further train VLMs using GRPO-lite with Backward Discounted Allocation (BDA), a reinforcement learning approach that removes KL regularization to encourage stronger policy optimization while allocating greater credit to the later reasoning steps where analytical conclusions are formed. Experiments on Qwen2.5-VL-7B and InternVL3-8B show that GRPO-lite fine-tuning on SD-MAR improves in-domain accuracy by up to 36.95%, with Qwen2.5-VL-7B outperforming GPT-4.1 on the SD-MAR benchmark. Importantly, out-of-domain generalization is preserved or improved: performance remains within 1% on MME, MMMU-Pro, and MathVista, while improving by up to 4% on MMBench. LLM-as-judge evaluation further demonstrates consistent improvements in logical coherence and explanation quality across both models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。