arXiv:2505.23265cs.CV2025-05

构建12.8万样本数据集,提升多模态模型的空间合理性推理能力

SPR-128K: A New Benchmark for Spatial Plausibility Reasoning with Multimodal Large Language Models

  • 构建SPR-128K数据集,覆盖四类空间合理性判断任务
  • 新奖励机制使小模型超越主流大模型的推理表现
  • 适合关注视觉推理与模型效率的研究者参考

近年来图像生成性能显著提升,但图像筛选研究较少,多模态大模型(MLLMs)在空间合理性推理方面表现不佳,主要受限于数据匮乏和模型本身的空间推理能力弱。本文从数据与方法两方面提出完整解决方案。数据层面,构建包含超过12.8万样本的时空合理性推理(SPR)数据集,名为SPR-128K,涵盖四个评估维度;在标注方面,探索多种高性价比获取高质量思维链(CoT)数据的方法。方法层面,将动态比例准确率(DPA)奖励引入组相对策略优化(GRPO)框架,提出DPA-GRPO。实验表明,即使领先大模型在空间合理性推理上表现仍不理想,而本研究所用的小模型通过DPA-GRPO显著超越多个开源及闭源主流模型。

原文摘要 · Abstract (English)

The performance of image generation has been significantly improved in recent years. However, the study of image screening is rare, and its performance with Multimodal Large Language Models (MLLMs) is unsatisfactory due to the lack of data and the weak spatial plausibility reasoning ability in MLLMs. In this work, we propose a complete solution to address these problems in terms of data and methodology. For data, we collect a comprehensive spatial plausibility reasoning (SPR) dataset with over 128k samples, called SPR-128K. The dataset evaluates spatial plausibility reasoning ability under four aspects. Regarding data annotation, we investigate multiple approaches to acquire high-quality Chain-of-Thought (CoT) data in the most cost-effective manner. Methodologically, we introduce a Dynamic Proportional Accuracy (DPA) reward into the Group Relative Policy Optimization (GRPO) framework, called DPA-GRPO. This enhanced method demonstrates superior performance compared to the original GRPO. Our experiments reveal that even leading MLLMs exhibit unsatisfactory performance in spatial plausibility reasoning. In contrast, our much smaller model, leveraging DPA-GRPO, substantially surpasses both large open-source and leading closed-source models.

多模态空间推理数据集强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。