arXiv:2409.04214cs.CV2024-09被引 17

让AI更懂几何图,提升数学推理能力

Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver

  • 用形式化语言+视觉特征融合增强几何理解
  • 在formalgeo7k上实现开放任务的准确率提升
  • 自建22.8万条带形式化标注的合成数据集

数学推理对AI仍是挑战,尤其涉及需结合语言与视觉信号的几何问题。由于多数多模态大模型(MLLMs)的视觉编码器基于自然场景训练,难以理解几何图,其解题表现仅略优于纯文本模型。这一局限源于缺乏有效的几何关系表征方法。为此,我们提出图示形式化增强几何求解框架(DFE-GPS),融合视觉特征、几何形式语言与自然语言表示。设计一种新型合成数据生成方法,构建大规模几何数据集SynthGeo228K,包含形式化与自然语言双标注,用于强化视觉编码器对几何结构的理解。该框架显著提升MLLM对几何图的处理能力,并拓展其在formalgeo7k数据集上开放任务的应用。

原文摘要 · Abstract (English)

Mathematical reasoning remains an ongoing challenge for AI models, especially for geometry problems that require both linguistic and visual signals. As the vision encoders of most MLLMs are trained on natural scenes, they often struggle to understand geometric diagrams, performing no better in geometry problem solving than LLMs that only process text. This limitation is amplified by the lack of effective methods for representing geometric relationships. To address these issues, we introduce the Diagram Formalization Enhanced Geometry Problem Solver (DFE-GPS), a new framework that integrates visual features, geometric formal language, and natural language representations. We propose a novel synthetic data approach and create a large-scale geometric dataset, SynthGeo228K, annotated with both formal and natural language captions, designed to enhance the vision encoder for a better understanding of geometric structures. Our framework improves MLLMs' ability to process geometric diagrams and extends their application to open-ended tasks on the formalgeo7k dataset.

几何推理多模态形式化语言合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。