arXiv:2412.11757cs.CL2024-12被引 13

构建涵盖多类型推理的科学表格与文本问答基准

SCITAT: A Question Answering Benchmark for Scientific Tables and Text Covering Diverse Reasoning Types

  • 融合表格与文本,覆盖真实科研场景中的多种推理类型
  • 新基线模型在该基准上平均提升12.9%性能
  • 适合研究科学问答、多模态推理的学者使用

科学问答(SQA)旨在基于论文回答问题。然而,现有SQA数据集推理类型有限,且忽视表格与文本间的关联,与真实科研场景存在显著差距。为此,我们提出一个涵盖多样化推理类型的科学表格与文本问答基准(SciTaT)。为覆盖更多推理类型,我们从真实问题中归纳出多种推理模式;为同时利用表格与文本,要求问题尽可能结合二者。基于SciTaT,我们提出一个强基线模型CaR,整合多种推理方法,同步处理表格与文本。CaR在SciTaT上相比其他基线平均提升12.9%,验证了其有效性。错误分析揭示了该基准的挑战,如复杂数值计算和领域知识依赖。

原文摘要 · Abstract (English)

Scientific question answering (SQA) is an important task aimed at answering questions based on papers. However, current SQA datasets have limited reasoning types and neglect the relevance between tables and text, creating a significant gap with real scenarios. To address these challenges, we propose a QA benchmark for scientific tables and text with diverse reasoning types (SciTaT). To cover more reasoning types, we summarize various reasoning types from real-world questions. To involve both tables and text, we require the questions to incorporate tables and text as much as possible. Based on SciTaT, we propose a strong baseline (CaR), which combines various reasoning methods to address different reasoning types and process tables and text at the same time. CaR brings average improvements of 12.9% over other baselines on SciTaT, validating its effectiveness. Error analysis reveals the challenges of SciTaT, such as complex numerical calculations and domain knowledge.

科学问答多模态推理数据集构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。