arXiv:2503.05543cs.CVcs.CL2025-03ICCV被引 15

利用图形信息解决几何题中的文字歧义,提升推理准确率

Pi-GPS: Enhancing Geometry Problem Solving by Unleashing the Power of Diagrammatic Information

  • 用图文联合模块解析文本歧义,结合几何规则校验输出
  • 在Geometry3K数据集上比现有方法提升近10%准确率
  • 适合需要高精度数学推理的智能教育系统使用

几何问题求解因在智能教育领域的应用潜力而受到越来越多关注。受文本常存在歧义而图形可澄清这一现象启发,本文提出Pi-GPS框架,通过挖掘图形信息来消除文本歧义,这是以往研究中被忽视的关键环节。具体而言,设计了一个包含修正器与验证器的微型模块:修正器利用多模态大模型(MLLMs)基于图形上下文对文本进行消歧;验证器则确保修正后的输出符合几何规则,减少模型幻觉。此外,还探索了在消歧后的形式语言基础上,使用大语言模型(LLMs)作为定理预测器的效果。实验证明,Pi-GPS超越当前最优模型,在Geometry3K数据集上较先前神经符号方法提升近10%。本工作强调了在多模态数学推理中消除文本歧义的重要性,这是制约性能的关键因素。

原文摘要 · Abstract (English)

Geometry problem solving has garnered increasing attention due to its potential applications in intelligent education field. Inspired by the observation that text often introduces ambiguities that diagrams can clarify, this paper presents Pi-GPS, a novel framework that unleashes the power of diagrammatic information to resolve textual ambiguities, an aspect largely overlooked in prior research. Specifically, we design a micro module comprising a rectifier and verifier: the rectifier employs MLLMs to disambiguate text based on the diagrammatic context, while the verifier ensures the rectified output adherence to geometric rules, mitigating model hallucinations. Additionally, we explore the impact of LLMs in theorem predictor based on the disambiguated formal language. Empirical results demonstrate that Pi-GPS surpasses state-of-the-art models, achieving a nearly 10\% improvement on Geometry3K over prior neural-symbolic approaches. We hope this work highlights the significance of resolving textual ambiguity in multimodal mathematical reasoning, a crucial factor limiting performance.

几何推理多模态消歧

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。