构建细粒度科学文献问答数据集,提升模型证据定位与推理能力
SciEGQA: A Dataset for Scientific Evidence-Grounded Question Answering and Reasoning
- 标注支持答案的语义连贯文本区域及边界框,实现证据精准定位
- 包含1623个高质量问答对,训练集超3万对,覆盖复杂科学文档
- 显著提升视觉语言模型在科学文献中的推理表现,适合研究多模态理解者
科学文献具有复杂的多模态结构,使文档视觉问答中的证据定位与科学推理尤为困难。然而,现有大多数基准仅在页面层面评估模型,未明确标注支撑答案的证据区域,限制了评估的可解释性与可靠性。为此,我们提出SciEGQA,一个具备语义证据接地的科学文档问答与推理数据集,其中支持性证据以带有边界框的语义连贯文档区域表示。SciEGQA包含两部分:一个由人工标注的细粒度基准集(1,623个高质量问答对),以及一个通过自动化数据构建流程生成的大规模训练集(超过30,000个QA对)。对多种视觉-语言模型(VLMs)的广泛实验表明,现有模型在科学文档中仍难以完成证据定位与基于证据的问答。在该数据集上训练能显著提升VLMs的科学推理能力。
原文摘要 · Abstract (English)
Scientific documents contain complex multimodal structures, which makes evidence localization and scientific reasoning in Document Visual Question Answering particularly challenging. However, most existing benchmarks evaluate models only at the page level without explicitly annotating the evidence regions that support the answer, which limits both interpretability and the reliability of evaluation. To address this limitation, we introduce SciEGQA, a scientific document question answering and reasoning dataset with semantic evidence grounding, where supporting evidence is represented as semantically coherent document regions annotated with bounding boxes. SciEGQA consists of two components: a **human-annotated fine-grained benchmark** containing 1,623 high-quality question--answer pairs, and a **large-scale automatically constructed training set** with over 30K QA pairs generated through an automated data construction pipeline. Extensive experiments on a wide range of Vision-Language Models (VLMs) show that existing models still struggle with evidence localization and evidence-based question answering in scientific documents. Training on the proposed dataset significantly improves the scientific reasoning capabilities of VLMs. The project page is available at https://yuwenhan07.github.io/SciEGQA-project/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。