arXiv:2511.23112cs.CVcs.LG2025-11ACL

测试视觉模型真能看懂数学题,发现越难的题越不靠图。

MathSight: A Benchmark Exploring Have Vision-Language Models Really Seen in University-Level Mathematical Reasoning?

  • 设计多版本数学题,对比有图、无图时模型表现差异。
  • 难题中图像贡献下降,无图模型反而优于带图版本。
  • 揭示当前模型依赖语言先验,适合评估视觉理解真实能力。

视觉-语言模型(VLMs)在多模态数学推理上取得显著进展,但其对视觉信息的真实利用程度仍不明确。现有基准报告了优异的整体性能,却很少分离图像模态的作用,难以判断模型是真正理解视觉内容,还是仅依赖语言先验。为此,我们提出MathSight,一个面向大学级别数学推理的多模态基准,旨在解耦并量化视觉输入的影响。每个题目包含原始图、手绘图、照片拍摄版及纯文本条件,实现受控对比。在主流VLMs上的实验显示:随着题目难度提升,视觉信息的贡献持续下降。值得注意的是,Qwen3-VL在无图像输入时的表现超越其多模态版本及GPT-5,凸显了构建类似MathSight的基准对于推动未来模型实现真正视觉感知推理的重要性。

原文摘要 · Abstract (English)

Recent advances in Vision-Language Models (VLMs) have achieved impressive progress in multimodal mathematical reasoning. Yet, how much visual information truly contributes to reasoning remains unclear. Existing benchmarks report strong overall performance but seldom isolate the role of the image modality, leaving open whether VLMs genuinely leverage visual understanding or merely depend on linguistic priors. To address this, we present MathSight, a university-level multimodal mathematical reasoning benchmark designed to disentangle and quantify the effect of visual input. Each problem includes multiple visual variants -- original, hand-drawn, photo-captured -- and a text-only condition for controlled comparison. Experiments on state-of-the-art VLMs reveal a consistent trend: the contribution of visual information diminishes with increasing problem difficulty. Remarkably, Qwen3-VL without any image input surpasses both its multimodal variants and GPT-5, underscoring the need for benchmarks like MathSight to advance genuine vision-grounded reasoning in future models.

多模态推理视觉理解数学模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。