用小模型+新方法,让AI更懂图像是否符合物理常识。
Physical Plausibility Reasoning via HCM-GRPO: Empowering Compact Model for Superior Performance
- 构建128k样本数据集,四类物理合理性评估,支持高效标注
- 提出HCM-GRPO框架,小模型超越大模型的物理推理能力
- 适合做图像生成质量评估或轻量化多模态系统研发
近年来图像生成性能显著提升,但图像筛选研究较少,且多模态大语言模型(MLLMs)在该任务上表现不佳,主要因数据稀缺和物理合理性推理能力弱。本文从数据与方法两方面提出完整解决方案:数据层面,构建包含超过12.8万样本、约64万张图像的图像筛选数据集,每样本含一张原始图与四张生成图,从外观变形、物理阴影、布局放置、扩展合理性四个维度评估物理合理性;通过人工、全自动及答案驱动等多重标注方式,在成本可控前提下获取高质量思维链(CoT)数据。方法层面,提出在组相对策略优化(GRPO)中引入硬样本挖掘(HCM)与动态比例准确率(DPA)奖励机制,形成HCM-GRPO。实验表明,即使是GPT5.2、Gemini3-Pro等先进闭源模型,在物理合理性推理上仍表现不佳;而采用HCM-GRPO后,仅用较小模型即可超越多个开源与主流闭源模型的表现。
原文摘要 · Abstract (English)
The performance of image generation has been significantly improved in recent years. However, the study of image screening is rare, and its performance with Multimodal Large Language Models (MLLMs) is unsatisfactory due to the lack of data and the weak physical plausibility reasoning ability in MLLMs. In this work, we propose a complete solution to address these problems in terms of data and methodology. For data, we collect a comprehensive image screening dataset with over 128k samples, comprising about 640k images. Each sample consists of an original image and four generated images. The dataset evaluates the physical plausibility reasoning ability under four aspects: appearance deformation, physical shadow, placement layout, and extension rationality. Regarding data annotation, we investigate multiple approaches, including purely manual, fully automated, and answer-driven annotations, to acquire high-quality chains of thought (CoT) data in the most cost-effective manner. Methodologically, we introduce a Hard Cases Mining (HCM) strategy with a Dynamic Proportional Accuracy (DPA) reward into the Group Relative Policy Optimization (GRPO) framework, called HCM-GRPO. This enhanced method demonstrates superior physical plausibility reasoning capabilities compared to the original GRPO. Our experimental results reveal that even state-of-the-art closed-source MLLMs, such as GPT5.2 and Gemini3-Pro, exhibit unsatisfactory performance in physical plausibility reasoning. In contrast, by leveraging the HCM-GRPO, we are able to surpass the scores of both large-scale open-source and leading closed-source models with a much smaller model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。