让大模型学会解立体几何题,还能自动生成解题步骤和图形描述。
Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning
- 用检索增强+微调训练,结合上下文学习提升解题能力。
- 在GeoQA和GeoMath数据集上分别达到65.25%和42.36%准确率。
- 首次支持立体几何求解,适合数学教育与多模态研究者。
几何数学问题对大语言模型(LLM)构成挑战,因其涉及视觉元素与空间推理。现有方法多依赖符号理解,但该领域数据稀缺,尤其缺乏立体几何题。为此,我们从中国高中教育网站收集数据,构建包含完整推理步骤的立体几何问答数据集GeoMath。提出名为Geo-LLaVA的大规模多模态模型框架,在训练阶段融合检索增强与监督微调(称为元训练),推理时采用上下文学习(ICL)以提升性能。微调后使用ICL的模型在GeoQA与GeoMath数据集上分别取得65.25%和42.36%的准确率。模型首次具备解决立体几何问题的能力,并能生成合理的立体图形描述与解题步骤。本研究为大模型在多模态数学问题求解,特别是几何领域的探索奠定基础。
原文摘要 · Abstract (English)
Geometry mathematics problems pose significant challenges for large language models (LLMs) because they involve visual elements and spatial reasoning. Current methods primarily rely on symbolic character awareness to address these problems. Considering geometry problem solving is a relatively nascent field with limited suitable datasets and currently almost no work on solid geometry problem solving, we collect a geometry question-answer dataset by sourcing geometric data from Chinese high school education websites, referred to as GeoMath. It contains solid geometry questions and answers with accurate reasoning steps as compensation for existing plane geometry datasets. Additionally, we propose a Large Multi-modal Model (LMM) framework named Geo-LLaVA, which incorporates retrieval augmentation with supervised fine-tuning (SFT) in the training stage, called meta-training, and employs in-context learning (ICL) during inference to improve performance. Our fine-tuned model with ICL attains the state-of-the-art performance of 65.25% and 42.36% on selected questions of the GeoQA dataset and GeoMath dataset respectively with proper inference steps. Notably, our model initially endows the ability to solve solid geometry problems and supports the generation of reasonable solid geometry picture descriptions and problem-solving steps. Our research sets the stage for further exploration of LLMs in multi-modal math problem-solving, particularly in geometry math problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。