arXiv:2506.12909cs.CL2025-06

用动态数值题评估大模型科学推理能力,避免记忆作弊。

SciDA: Scientific Dynamic Assessor of LLMs

  • 构建多学科动态数值题集,每次推理随机初始化数字
  • 顶级大模型在动态题上表现大幅下降,暴露记忆依赖问题
  • 适合评估模型真实推理能力,尤其关注数值计算的可信度

大语言模型(LLMs)推理能力的提升使其能更高效地解决科学问题。因此,一个高质量、全面且恰当的评估基准至关重要。然而,现有基准或存在数据污染风险,或缺乏跨学科覆盖。具体而言,由于训练数据与静态基准间存在数据源重叠,模型可能无意中记忆答案的关键特征或数字模式(即数据污染),导致对推理能力的系统性高估,尤其是数值推理。为此,我们提出SciDA,一个涵盖1000+项奥运级数值计算题的多学科基准,通过每轮推理随机初始化数值,避免依赖固定数字模式。我们对闭源和开源的顶尖大模型进行了系列实验,发现其在动态数值初始化下性能显著下降。这为大模型的数值推理能力提供了真实、无偏的评估。数据已公开于https://huggingface.co/datasets/m-a-p/SciDA。

原文摘要 · Abstract (English)

Advancement in Large Language Models (LLMs) reasoning capabilities enables them to solve scientific problems with enhanced efficacy. Thereby, a high-quality benchmark for comprehensive and appropriate assessment holds significance, while existing ones either confront the risk of data contamination or lack involved disciplines. To be specific, due to the data source overlap of LLMs training and static benchmark, the keys or number pattern of answers inadvertently memorized (i.e. data contamination), leading to systematic overestimation of their reasoning capabilities, especially numerical reasoning. We propose SciDA, a multidisciplinary benchmark that consists exclusively of over 1k Olympic-level numerical computation problems, allowing randomized numerical initializations for each inference round to avoid reliance on fixed numerical patterns. We conduct a series of experiments with both closed-source and open-source top-performing LLMs, and it is observed that the performance of LLMs drop significantly under random numerical initialization. Thus, we provide truthful and unbiased assessments of the numerical reasoning capabilities of LLMs. The data is available at https://huggingface.co/datasets/m-a-p/SciDA

大模型评估数值推理动态测试数据污染

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。