构建大规模科学多模态文档推理数据集,提升模型跨文档理解能力
SciMDR: Advancing Scientific Multimodal Document Reasoning
- 分两阶段生成:先聚焦命题生成问答对,再嵌入完整论文增强真实复杂性
- 构建30万条带推理链的问答对,覆盖2万篇科学论文,支持复杂推理任务
- 适合训练需要长文档理解的科学智能模型,尤其擅长多步逻辑推理
为解决科学多模态文档推理数据集在规模、忠实度与真实性之间的权衡问题,本文提出合成-重锚框架(synthesize-and-reground),包含两个阶段:(1) 命题中心型QA合成,生成忠实且独立的问答对及针对特定段落的推理;(2) 文档级重锚,将这些问答对程序化地嵌入完整文档任务中,确保真实复杂性。基于此框架,我们构建了SciMDR——一个用于跨模态理解的大规模训练数据集,包含30万条问答对,涵盖2万篇科学论文,并带有明确的推理链。同时构建SciMDR-Eval,作为专家标注的基准测试集,用于评估完整科学流程中的多模态理解能力。实验表明,基于SciMDR微调的模型在多个科学问答基准上表现显著提升,尤其在需要复杂文档级推理的任务中优势明显。
原文摘要 · Abstract (English)
Constructing scientific multimodal document reasoning datasets for foundation model training involves an inherent trade-off among scale, faithfulness, and realism. To address this challenge, we introduce the synthesize-and-reground framework, a two-stage pipeline comprising: (1) Claim-Centric QA Synthesis, which generates faithful, isolated QA pairs and reasoning on focused segments, and (2) Document-Scale Regrounding, which programmatically re-embeds these pairs into full-document tasks to ensure realistic complexity. Using this framework, we construct SciMDR, a large-scale training dataset for cross-modal comprehension, comprising 300K QA pairs with explicit reasoning chains across 20K scientific papers. We further construct SciMDR-Eval, an expert-annotated benchmark to evaluate multimodal comprehension within full-length scientific workflows. Experiments demonstrate that models fine-tuned on SciMDR achieve significant improvements across multiple scientific QA benchmarks, particularly in those tasks requiring complex document-level reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。