构建科学领域评测数据集,评估大模型理解复杂科研内容的能力
SciCUEval: A Comprehensive Dataset for Evaluating Scientific Context Understanding in Large Language Models
- 设计十类学科子数据集,覆盖生物到材料科学的多模态科研数据
- 从信息识别到跨源推理,系统测试大模型在五个科学任务中的表现
- 适合关注科研AI模型能力评估的研究者和开发者参考
大型语言模型在上下文理解与推理方面展现出惊人能力,但现有评估基准多聚焦通用领域,难以捕捉科学数据的复杂性。为此,我们构建了面向科学上下文理解能力评估的综合基准数据集SciCUEval,包含涵盖生物学、化学、物理学、生物医学和材料科学的十个领域特定子数据集,整合结构化表格、知识图谱与非结构化文本等多种数据模态。该数据集系统评估四项核心能力:相关信息识别、信息缺失检测、多源信息融合与上下文感知推理,并通过多种题型实现全面评测。我们对前沿大模型在SciCUEval上的表现进行了广泛评估,提供了其在科学上下文理解方面的优劣势细粒度分析,为未来科学领域大模型的发展提供重要参考。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown impressive capabilities in contextual understanding and reasoning. However, evaluating their performance across diverse scientific domains remains underexplored, as existing benchmarks primarily focus on general domains and fail to capture the intricate complexity of scientific data. To bridge this gap, we construct SciCUEval, a comprehensive benchmark dataset tailored to assess the scientific context understanding capability of LLMs. It comprises ten domain-specific sub-datasets spanning biology, chemistry, physics, biomedicine, and materials science, integrating diverse data modalities including structured tables, knowledge graphs, and unstructured texts. SciCUEval systematically evaluates four core competencies: Relevant information identification, Information-absence detection, Multi-source information integration, and Context-aware inference, through a variety of question formats. We conduct extensive evaluations of state-of-the-art LLMs on SciCUEval, providing a fine-grained analysis of their strengths and limitations in scientific context understanding, and offering valuable insights for the future development of scientific-domain LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。