测试文字转图像模型在数学可视化任务上的表现,发现准确率普遍极低。
MathGen: Revealing the Illusion of Mathematical Competence through Text-to-Image Generation
- 构建420道数学视觉生成题的基准,分清洁与开放场景两类。
- 顶尖闭源模型准确率仅53.7%,开源模型普遍1%~11%。
- 对几何和函数绘图要求精确的题目,模型几乎无法达标。
现代生成模型虽能解决复杂数学问题,但在真实场景中,数学解答常需通过图表、函数图像、几何构造等视觉形式表达,其正确性依赖精准的视觉布局。这引出关键问题:当答案必须以图像呈现而非文本时,生成模型是否仍能胜任?为此,我们提出MathGen,一个包含420道题的严格基准,涵盖七个核心领域,包括350个Clean-Scene问题和70个Open-Scene配对问题。每个问题采用Script-as-a-Judge协议,通过可复用的可执行脚本实现问题特定的验证逻辑,确保评估的确定性与可复现性。在代表性开源与专有文字转图像模型上的实验表明,数学保真度仍是主要瓶颈:即使最优闭源模型整体准确率也仅为53.7%,而开源模型仅达1%–11%,在需要精确几何与函数渲染的结构化任务上常接近0%。总体而言,当前T2I模型在基础数学视觉生成任务上仍不可靠。
原文摘要 · Abstract (English)
Modern generative models have demonstrated the ability to solve challenging mathematical problems. In many real-world settings, however, mathematical solutions must be expressed visually through diagrams, plots, geometric constructions, and structured symbolic layouts, where correctness depends on precise visual composition. This naturally raises the question of whether generative models can still do so when the answer must be rendered visually rather than written in text? To study this problem, we introduce MathGen, a rigorous benchmark of 420 problems spanning seven core domains, including 350 Clean-Scene problems and 70 paired Open-Scene problems. Each problem is evaluated under a Script-as-a-Judge protocol with problem-specific verification criteria implemented through reusable executable scripts for deterministic and reproducible evaluation. Experiments on representative open-source and proprietary text-to-image models show that mathematical fidelity remains a major bottleneck: even the best closed-source model reaches only 53.7% overall accuracy, while open-source models achieve just 1--11%, often near 0% on structured tasks, particularly those requiring precise geometric and functional rendering. Overall, current T2I models remain far from reliable at even elementary mathematical visual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。