arXiv:2506.06034cs.CL2025-06被引 12

构建首个多模态定理证明基准,测试大模型能否像人一样看图证数。

MATP-BENCH: Can MLLM Be a Good Automated Theorem Prover for Multimodal Problems?

  • 设计包含1056道多模态数学题的基准,融合图形与形式化语言。
  • 现有大模型仅能解决少数题目,说明多模态自动证明仍存巨大挑战。
  • 适合研究多模态推理、形式化证明和数学AI的学者参考。

众多定理(如几何)常以图文混合形式呈现。人类在该场景下可通过视觉推理获得直觉并指导证明过程。现代多模态大语言模型(MLLMs)在解决各类数学问题上表现卓越,但其作为自动化定理证明器(ATP)在多模态领域的潜力尚未被充分探索。本文提出多模态自动化定理证明基准(MATP-BENCH),这是一个多模态、多层级、多语言的基准,用于评估MLLMs在该任务中的表现。该基准包含1056道来自中学、大学及竞赛水平的数学题,每道题均配有Lean 4、Coq和Isabelle等形式化表达,兼容多种定理证明框架。该任务要求模型整合复杂的视觉理解、广泛数学知识与严格的符号推理能力,生成形式化证明。我们使用MATP-BENCH评估多种先进多模态语言模型,结果表明现有方法仅能解决有限题目,凸显该基准对自动定理证明研究的开放性挑战。

原文摘要 · Abstract (English)

Numerous theorems, such as those in geometry, are often presented in multimodal forms (e.g., diagrams). Humans benefit from visual reasoning in such settings, using diagrams to gain intuition and guide the proof process. Modern Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities in solving a wide range of mathematical problems. However, the potential of MLLMs as Automated Theorem Provers (ATPs), specifically in the multimodal domain, remains underexplored. In this paper, we introduce the Multimodal Automated Theorem Proving benchmark (MATP-BENCH), a new Multimodal, Multi-level, and Multi-language benchmark designed to evaluate MLLMs in this role as multimodal automated theorem provers. MATP-BENCH consists of 1056 multimodal theorems drawn from high school, university, and competition-level mathematics. All these multimodal problems are accompanied by formalizations in Lean 4, Coq and Isabelle, thus making the benchmark compatible with a wide range of theorem-proving frameworks. MATP-BENCH requires models to integrate sophisticated visual understanding with mastery of a broad spectrum of mathematical knowledge and rigorous symbolic reasoning to generate formal proofs. We use MATP-BENCH to evaluate a variety of advanced multimodal language models. Existing methods can only solve a limited number of the MATP-BENCH problems, indicating that this benchmark poses an open challenge for research on automated theorem proving.

定理证明多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。