新基准挑战大模型科学方程发现能力,真实评估其推理水平。
LLM-SRBench: A New Benchmark for Scientific Equation Discovery with Large Language Models
- 设计两类难题:变换公式形式与合成数据驱动问题
- 顶尖模型符号准确率仅31.5%,暴露当前能力瓶颈
- 适合研究科学推理、大模型泛化能力的学者使用
科学方程发现是科学进步的核心任务,可推导自然现象的规律。近年来,大型语言模型(LLMs)因其蕴含的科学知识潜力,被用于生成假设。然而,现有评估基准多依赖常见方程,易被模型记忆,导致性能指标虚高,无法反映真实发现能力。本文提出LLM-SRBench,一个涵盖四个科学领域的239个挑战性问题的综合性基准,专为评估基于LLM的科学方程发现方法而设计,防止简单记忆。该基准包含两类:LSR-Transform将常见物理模型转换为不常见的数学表达形式,检验超越记忆的推理能力;LSR-Synth引入合成的、以发现为导向的问题,要求数据驱动推理。对多个前沿方法(包括开源与闭源LLM)的广泛评估显示,目前最佳系统符号准确率仅为31.5%。这些结果凸显了科学方程发现的难度,确立了LLM-SRBench作为未来研究的重要资源。
原文摘要 · Abstract (English)
Scientific equation discovery is a fundamental task in the history of scientific progress, enabling the derivation of laws governing natural phenomena. Recently, Large Language Models (LLMs) have gained interest for this task due to their potential to leverage embedded scientific knowledge for hypothesis generation. However, evaluating the true discovery capabilities of these methods remains challenging, as existing benchmarks often rely on common equations that are susceptible to memorization by LLMs, leading to inflated performance metrics that do not reflect discovery. In this paper, we introduce LLM-SRBench, a comprehensive benchmark with 239 challenging problems across four scientific domains specifically designed to evaluate LLM-based scientific equation discovery methods while preventing trivial memorization. Our benchmark comprises two main categories: LSR-Transform, which transforms common physical models into less common mathematical representations to test reasoning beyond memorized forms, and LSR-Synth, which introduces synthetic, discovery-driven problems requiring data-driven reasoning. Through extensive evaluation of several state-of-the-art methods, using both open and closed LLMs, we find that the best-performing system so far achieves only 31.5% symbolic accuracy. These findings highlight the challenges of scientific equation discovery, positioning LLM-SRBench as a valuable resource for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。