填补非计算机领域科学文本推理空白,构建首个跨学科评测集
A MISMATCHED Benchmark for Scientific Natural Language Inference
- 构建涵盖心理、工程、公共健康三领域的科学文本推理数据集
- 2700对人工标注句对,最佳模型宏平均F1仅78.17%
- 首次证明隐含科学推理关系有助于提升模型性能
科学自然语言推理(NLI)旨在判断科研论文中两句话之间的语义关系。现有数据集集中于计算机科学领域,而其他非计算机领域完全被忽视。本文提出新型科学NLI评测基准MISMATCHED,覆盖心理学、工程学和公共卫生三个非计算机领域,包含2700对人工标注的句子对。我们基于预训练小语言模型(SLMs)和大语言模型(LLMs)建立了强基线,最优模型的宏平均F1仅为78.17%,表明仍有巨大提升空间。此外,研究发现将具有隐含科学推理关系的句对纳入训练可有效提升模型表现。数据集与代码已开源至GitHub。
原文摘要 · Abstract (English)
Scientific Natural Language Inference (NLI) is the task of predicting the semantic relation between a pair of sentences extracted from research articles. Existing datasets for this task are derived from various computer science (CS) domains, whereas non-CS domains are completely ignored. In this paper, we introduce a novel evaluation benchmark for scientific NLI, called MISMATCHED. The new MISMATCHED benchmark covers three non-CS domains-PSYCHOLOGY, ENGINEERING, and PUBLIC HEALTH, and contains 2,700 human annotated sentence pairs. We establish strong baselines on MISMATCHED using both Pre-trained Small Language Models (SLMs) and Large Language Models (LLMs). Our best performing baseline shows a Macro F1 of only 78.17% illustrating the substantial headroom for future improvements. In addition to introducing the MISMATCHED benchmark, we show that incorporating sentence pairs having an implicit scientific NLI relation between them in model training improves their performance on scientific NLI. We make our dataset and code publicly available on GitHub.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。