构建1.5万道可执行代码生成答案的物理题库,测试模型科学推理能力
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
- 用参数化问题+可执行代码生成答案,支持无限变体
- 提出一致性、失败率、混淆率新指标,评估模型稳定性
- 适合研究模型在复杂科学推理中的鲁棒性与可解释性
我们提出一个大规模合成基准,包含15,045道大学水平物理题(训练/测试集比例90:10)。每道题完全参数化,支持近乎无限的输入配置,并附有结构化逐步推理和可执行的Python代码,能为任意参数组合生成真值解。基准涵盖三种题型:MC-Symbolic(符号选项多选题)、MC-Numerical(数值选项多选题)和自由作答题,测试互补的推理能力。利用动态代码驱动特性,我们引入三项新评估指标:一致性得分、失败率和混淆率,量化不同题型变体下的结果变异性和不确定性。对先进指令微调语言模型的实验揭示了其在科学推理中的优势与局限,确立SymPyBench作为构建更鲁棒、可解释推理系统的基础。
原文摘要 · Abstract (English)
We introduce, a large-scale synthetic benchmark of 15,045 university-level physics problems (90/10% train/test split). Each problem is fully parameterized, supporting an effectively infinite range of input configurations, and is accompanied by structured, step-by-step reasoning and executable Python code that produces the ground-truth solution for any parameter set. The benchmark contains three question types: MC-Symbolic (multiple-choice with symbolic options), MC-Numerical (multiple-choice with numerical options), and free-form (open-ended responses). These diverse formats test complementary reasoning skills. By leveraging the dynamic, code-driven nature of the benchmark, we introduce three novel evaluation metrics in addition to standard accuracy: Consistency Score, Failure Rate, and Confusion Rate, that quantify variability and uncertainty across problem variants. Experiments with state-of-the-art instruction-tuned language models reveal both strengths and limitations in scientific reasoning, positioning SymPyBench as a foundation for developing more robust and interpretable reasoning systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。