用强化学习生成几何图像描述,提升模型泛化能力
Generalizable Geometric Image Caption Synthesis
- 引入可验证奖励的强化学习优化几何图像描述生成
- 在数学与工程任务中实现2.8%~4.8%准确率提升
- 适合需要几何推理的多模态模型训练场景
多模态大模型虽有广泛应用,但在复杂几何问题上仍表现不足。主要瓶颈在于缺乏高质量的图像-文本配对数据集,且传统模板生成方法难以泛化。本文提出将可验证奖励的强化学习(RLVR)融入数据生成流程,基于50种基本几何关系合成图像,并利用数学求解任务的奖励信号优化描述。该方法有效捕捉几何问题求解的关键特征,显著提升模型泛化能力。即使在分布外场景下,生成的数据集仍使多模态大模型在MathVista和MathVerse的统计、算术、代数与数值任务中准确率提升2.8%~4.8%,在MMMU的美术、设计、技术与工程任务中提升2.4%~3.9%。
原文摘要 · Abstract (English)
Multimodal large language models have various practical applications that demand strong reasoning abilities. Despite recent advancements, these models still struggle to solve complex geometric problems. A key challenge stems from the lack of high-quality image-text pair datasets for understanding geometric images. Furthermore, most template-based data synthesis pipelines typically fail to generalize to questions beyond their predefined templates. In this paper, we bridge this gap by introducing a complementary process of Reinforcement Learning with Verifiable Rewards (RLVR) into the data generation pipeline. By adopting RLVR to refine captions for geometric images synthesized from 50 basic geometric relations and using reward signals derived from mathematical problem-solving tasks, our pipeline successfully captures the key features of geometry problem-solving. This enables better task generalization and yields non-trivial improvements. Furthermore, even in out-of-distribution scenarios, the generated dataset enhances the general reasoning capabilities of multimodal large language models, yielding accuracy improvements of $2.8\%\text{-}4.8\%$ in statistics, arithmetic, algebraic, and numerical tasks with non-geometric input images of MathVista and MathVerse, along with $2.4\%\text{-}3.9\%$ improvements in Art, Design, Tech, and Engineering tasks in MMMU.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。