arXiv:2412.00846cs.AI2024-12被引 5

构建几何多模态数据集,提升大模型解题与推理能力

Improving Multimodal LLMs Ability In Geometry Problem Solving, Reasoning, And Multistep Scoring

  • 构建5340个带步骤解答的几何题目,支持多步推理评估
  • 微调后模型在几何解题准确率上显著提升,验证数据有效性
  • 适合研究视觉语言模型几何推理能力的学者使用

本文提出GPSM4K,一个面向几何问题求解、推理与多步评分的综合性多模态数据集。该数据集包含从7至12年级数学教材中手工提取的2157组图文问答对,并扩充至5340个问题,涵盖数值计算与定理证明两类题型。与仅含选择题的PGPS9k、Geometry3K和Geo170K不同,GPSM4K提供一致格式的详细分步解答,可全面评估模型求解过程。测试集评估表明,开源语言模型在几何求解方面仍有提升空间。在训练集上微调可有效增强模型能力。此外,我们评估了图像描述与检索增强生成(RAG)对性能的影响。通过LLM自动比对真实答案与预测解法,实现最终答案的自动化评判。本研究有助于评估并提升LVLM的几何推理能力。

原文摘要 · Abstract (English)

This paper presents GPSM4K, a comprehensive geometry multimodal dataset tailored to augment the problem-solving capabilities of Large Vision Language Models (LVLMs). GPSM4K encompasses 2157 multimodal question-answer pairs manually extracted from mathematics textbooks spanning grades 7-12 and is further augmented to 5340 problems, consisting of both numerical and theorem-proving questions. In contrast to PGPS9k, Geometry3K, and Geo170K which feature only objective-type questions, GPSM4K offers detailed step-by-step solutions in a consistent format, facilitating a comprehensive evaluation of problem-solving approaches. This dataset serves as an excellent benchmark for assessing the geometric reasoning capabilities of LVLMs. Evaluation of our test set shows that there is scope for improvement needed in open-source language models in geometry problem-solving. Finetuning on our training set increases the geometry problem-solving capabilities of models. Further, We also evaluate the effectiveness of techniques such as image captioning and Retrieval Augmentation generation (RAG) on model performance. We leveraged LLM to automate the task of final answer evaluation by providing ground truth and predicted solutions. This research will help to assess and improve the geometric reasoning capabilities of LVLMs.

几何推理多模态数据集大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。