arXiv:2512.05930cs.AI2025-12

用可执行代码评估科学推理,揭示模型真实能力

PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation

  • 基于动态生成的物理数学题和可运行的Python代码验证
  • 涵盖2.47万道题目,支持细粒度错误分析与鲁棒性测试
  • 适合研究多模态模型在科学推理中的缺陷与改进

评估视觉语言模型(VLMs)在数学与物理等科学领域的表现面临独特挑战,不仅需预测最终答案,更需理解概念、进行符号推理并遵守形式化规律。现有基准普遍静态、缺乏中间推理步骤、对变化不鲁棒,且缺少科学正确性的验证机制。为此,我们提出PRiSM,一个基于可执行Python代码的合成、完全动态的多模态科学推理基准。PRiSM包含超过24,750道大学级物理与数学问题,通过可扩展的代理式流程PrismAgent生成结构化问题实例。每个问题均含动态文本与视觉输入、生成图像,以及丰富结构化输出:用于生成与验证真值的可执行Python代码和详细分步推理过程。其动态特性与基于Python的自动化真值生成,使我们能对多模态VLM进行细粒度实验审计,揭示其失败模式、不确定性行为及科学推理局限性。我们设计了五项针对性评估任务,涵盖泛化、符号程序合成、扰动鲁棒性、推理纠错与模糊性解析。通过对现有VLM的全面评估,我们揭示其不足,并展示PRiSM如何深入洞察模型的科学推理能力。

原文摘要 · Abstract (English)

Evaluating vision-language models (VLMs) in scientific domains like mathematics and physics poses unique challenges that go far beyond predicting final answers. These domains demand conceptual understanding, symbolic reasoning, and adherence to formal laws, requirements that most existing benchmarks fail to address. In particular, current datasets tend to be static, lacking intermediate reasoning steps, robustness to variations, or mechanisms for verifying scientific correctness. To address these limitations, we introduce PRiSM, a synthetic, fully dynamic, and multimodal benchmark for evaluating scientific reasoning via grounded Python code. PRiSM includes over 24,750 university-level physics and math problems, and it leverages our scalable agent-based pipeline, PrismAgent, to generate well-structured problem instances. Each problem contains dynamic textual and visual input, a generated figure, alongside rich structured outputs: executable Python code for ground truth generation and verification, and detailed step-by-step reasoning. The dynamic nature and Python-powered automated ground truth generation of our benchmark allow for fine-grained experimental auditing of multimodal VLMs, revealing failure modes, uncertainty behaviors, and limitations in scientific reasoning. To this end, we propose five targeted evaluation tasks covering generalization, symbolic program synthesis, perturbation robustness, reasoning correction, and ambiguity resolution. Through comprehensive evaluation of existing VLMs, we highlight their limitations and showcase how PRiSM enables deeper insights into their scientific reasoning capabilities.

科学推理多模态评估Python验证动态基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。