arXiv:2411.05338cs.CL2024-11EMNLP被引 24

构建首个由专家评审与作者作答的科学论文深度阅读理解数据集

SciDQA: A Deep Reading Comprehension Dataset over Scientific Papers

  • 基于领域专家评审和作者回答构建高质量科学问答对
  • 2,937组问题需跨图表、公式、附录及多文档推理
  • 适合评估大模型在复杂科学文本理解中的真实能力

科学文献通常信息密集,需深厚背景知识与深度理解才能有效阅读。我们提出SciDQA,一个面向科学论文深度阅读理解的新数据集,包含2,937个问答对。与现有科学QA数据集不同,SciDQA的问题源自领域专家的同行评审,答案由论文作者提供,确保对文献的全面检验。通过精细筛选低质量问题、去上下文化处理、追踪文档版本变化,并引入参考文献以支持多文档问答,提升了数据集质量。问题设计要求跨图表、表格、公式、附录及补充材料进行推理,且需多文档关联分析。我们对多种开源与专有大模型在不同配置下进行评估,使用表面相似度指标与大模型判断进行综合评测,揭示了显著性能差异。SciDQA是一个严格筛选、自然生成的科学问答数据集,旨在推动复杂科学文本理解的研究。

原文摘要 · Abstract (English)

Scientific literature is typically dense, requiring significant background knowledge and deep comprehension for effective engagement. We introduce SciDQA, a new dataset for reading comprehension that challenges LLMs for a deep understanding of scientific articles, consisting of 2,937 QA pairs. Unlike other scientific QA datasets, SciDQA sources questions from peer reviews by domain experts and answers by paper authors, ensuring a thorough examination of the literature. We enhance the dataset's quality through a process that carefully filters out lower quality questions, decontextualizes the content, tracks the source document across different versions, and incorporates a bibliography for multi-document question-answering. Questions in SciDQA necessitate reasoning across figures, tables, equations, appendices, and supplementary materials, and require multi-document reasoning. We evaluate several open-source and proprietary LLMs across various configurations to explore their capabilities in generating relevant and factual responses. Our comprehensive evaluation, based on metrics for surface-level similarity and LLM judgements, highlights notable performance discrepancies. SciDQA represents a rigorously curated, naturally derived scientific QA dataset, designed to facilitate research on complex scientific text understanding.

阅读理解科学文本大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。