arXiv:2510.21881cs.AIcs.CL2025-10被引 1

构建几何推理数据集,提升视觉语言模型的数学几何解题能力。

GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models

  • 构建包含6243个样本的GeoThought数据集,含逐步推理链和反思步骤。
  • 新模型在几何任务上超越基准,跨域泛化能力显著增强。
  • 适合研究多模态数学推理与思维链训练的研究者使用。

大型语言模型在文本数学问题求解中表现优异,但应用于视觉推理任务(尤其是几何题)时性能大幅下降,主要因几何问题需精细图像理解与多步推理,且现有数据集规模小、多样性不足、缺乏显式推理路径。为此,我们构建了GeoThought数据集,包含两个子集:Geo-Thought-6K(6,243个样本)及其增强版Geo-Thought-Augmented-10K(10,834个样本)。每条数据包含视觉描述、分步解答、明确推理链、反思步骤和最终答案。基于该数据集,我们开发了GeoThought-MLLM,一种能生成详细思维过程的多模态数学推理模型。实验表明,该模型在几何任务上优于现有基准,且在域内与域外设置中均表现更优。分析失败案例发现,错误主要源于数学概念误读或空间判断失误;通过引入思维链(CoT)修正后,模型可生成正确答案。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated strong reasoning capabilities in text-based mathematical problem solving; however, when adapted to visual reasoning tasks, particularly geometric problem solving, their performance substantially declines because geometric problems present unique challenges. Specifically, these challenges stem from two key factors: first, the intrinsic complexity of geometry requiring detailed image comprehension and multi-step reasoning, and second, the limitations of existing datasets which lack sufficient scale, diversity, and explicit reasoning traces, consequently hindering effective model training. To address these challenges, we developed the GeoThoughts dataset, a comprehensive geometric reasoning corpus with two subsets: Geo-Thought-6K with 6,243 samples and its augmented version Geo-Thought-Augmented-10K containing 10,834 samples. Each entry includes visual descriptions, step-by-step solutions, explicit reasoning chains, reflection steps, and final answers. Using this dataset, we developed GeoThought-MLLM, a mathematical reasoning multimodal model that generates detailed thinking processes during problem-solving. Our model outperforms existing benchmarks in geometric tasks, demonstrating that training with our Chain-of-Thought dataset improves geometric reasoning capabilities across both in-domain and out-of-domain settings. Finally, we analyze failure cases and observe that errors primarily arise from incorrect interpretation of mathematical concepts or spatial misjudgment. By invoking CoT to correct these mistakes, the model produces correct answers.

几何推理多模态思维链数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。