构建新材料科学可行性评估基准,避免大模型训练数据污染
SFBench: The SciFy Scientific Feasibility Benchmark

- 自建197个新材料领域科学主张,专家标注可行性评分与解释
- 采用五分制评分标准,要求开放性理由生成而非选择题
- 专为测试大模型科学推理能力设计,适合研究可信AI的学者
我们提出SFBench,一个用于评估系统判断科学主张可行性的基准数据集。该数据集包含197个材料科学领域的科学主张,每个主张均配有由领域专家评定的五分制可行性评分及详细解释。与以往数据集不同,SFBench具有四大特点:第一,任务复杂,需对多样可行性水平的主张进行推理;第二,主张为全新构建,非来自已有文献,显著降低大模型训练数据污染风险;第三,评分与标签由领域专家确立,而非人工智能生成;第四,评估不采用问答、多选或固定答案形式,解释要求完全开放。本文介绍基准设计、数据创建流程和评估指标,并报告使用近期GPT模型的基线结果。
原文摘要 · Abstract (English)
We present SFBench, a benchmark dataset for evaluating systems that assess the feasibility of scientific claims. SFBench includes 197 claims in materials science, each annotated with a ground-truth feasibility score on a five-point scale along with an explanation of that assessment. The collection differs from previous collections in several important ways: 1) it defines a complex task that requires reasoning over claims of varying scientific feasibility; 2) its claims are not extracted from existing scientific publications but are created de novo, greatly reducing the chances that LLMs have trained on them; 3) claims and ground truth are established by subject matter experts, not by artificial intelligence; and 4) unlike many benchmarks that ask about question/answer pairs, provide multiple choice answers, or ask questions requiring short, fixed answers, SFBench explanations are completely open-ended. We describe the benchmark design, data creation process, and evaluation metrics, and we report baseline results using recent GPT models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。