arXiv:2412.11863cs.CVcs.CL2024-12ICLR被引 39

GeoX通过统一视觉语言预训练提升几何问题求解能力。

GeoX: Geometric Problem Solving Through Unified Formalized Vision-Language Pre-training

  • 构建专用图符编码器与符号解码器,强化几何图像理解
  • 在多个基准上超越通用模型与专用模型,最高提升12.3%
  • 适合需要严谨几何推理的教育、竞赛场景

尽管多模态大语言模型在通用任务中表现优异,但在几何问题求解(GPS)上仍面临挑战,需理解图表、解析符号并进行复杂推理。这源于其在自然图像与文本上的预训练,以及缺乏自动验证机制。现有几何专用模型则受限于任务特定设计,泛化能力不足。为此,我们提出GeoX,一种专注于几何理解与推理的多模态大模型。针对几何图符与自然图像-文本的本质差异,引入单模态预训练,构建图符编码器与符号解码器,增强对几何图像与语料的理解。进一步提出几何-语言对齐预训练范式,弥合单模态专家间的模态鸿沟。设计生成-采样变压器(GS-Former),生成判别性查询,消除分布不均信号中的冗余表示。最后,通过视觉指令微调,使模型能以几何图像与问题为输入,输出可验证的解。实验表明,GeoX在公开基准如GeoQA、UniGeo、Geometry3K和PGPS9k上均优于通用模型与专用模型。

原文摘要 · Abstract (English)

Despite their proficiency in general tasks, Multi-modal Large Language Models (MLLMs) struggle with automatic Geometry Problem Solving (GPS), which demands understanding diagrams, interpreting symbols, and performing complex reasoning. This limitation arises from their pre-training on natural images and texts, along with the lack of automated verification in the problem-solving process. Besides, current geometric specialists are limited by their task-specific designs, making them less effective for broader geometric problems. To this end, we present GeoX, a multi-modal large model focusing on geometric understanding and reasoning tasks. Given the significant differences between geometric diagram-symbol and natural image-text, we introduce unimodal pre-training to develop a diagram encoder and symbol decoder, enhancing the understanding of geometric images and corpora. Furthermore, we introduce geometry-language alignment, an effective pre-training paradigm that bridges the modality gap between unimodal geometric experts. We propose a Generator-And-Sampler Transformer (GS-Former) to generate discriminative queries and eliminate uninformative representations from unevenly distributed geometric signals. Finally, GeoX benefits from visual instruction tuning, empowering it to take geometric images and questions as input and generate verifiable solutions. Experiments show that GeoX outperforms both generalists and geometric specialists on publicly recognized benchmarks, such as GeoQA, UniGeo, Geometry3K, and PGPS9k.

几何推理多模态视觉语言预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。