arXiv:2603.21193cs.CLcs.AI2026-03

从论文全文中精准定位假设与证据,提升科研发现整合效率。

Context Selection for Hypothesis and Statistical Evidence Extraction from Full-Text Scientific Articles

  • 分阶段检索+抽取框架,优化上下文选择以提升准确性。
  • 上下文质量优于数量,清理无关内容可显著提高效果。
  • 证据抽取仍困难,核心瓶颈在模型理解混合文本数据能力。

从完整科学论文中提取假设及其支持的统计证据,是整合实证发现的关键,但因文档长且论证分布于不同章节而极具挑战。本文研究一种序列式全文提取任务:将论文摘要中的主要发现与正文中的对应假设及支撑证据关联。该任务构成复杂的文档内检索问题,因大量候选段落主题相关但修辞角色不同,形成难样本。采用两阶段检索-抽取框架,系统评估检索设计选择,包括上下文数量、质量(标准RAG、重排序、微调检索器+重排序)以及理想段落设定,以区分检索失败与抽取限制。实验使用四种大语言模型抽取器,在四个数据集上验证。结果表明,针对性上下文选择显著优于全文本提示,增益集中于提升检索质量与上下文纯净度的配置;而统计证据抽取依然困难,即使使用理想段落,性能仍中等,表明根本瓶颈在于模型处理数值-文本混合表述的能力,而非检索本身。

原文摘要 · Abstract (English)

Extracting hypotheses and their supporting statistical evidence from full-text scientific articles is central to the synthesis of empirical findings, but remains difficult due to document length and the distribution of scientific arguments across sections of the paper. The work studies a sequential full-text extraction setting, where the statement of a primary finding in an article's abstract is linked to (i) a corresponding hypothesis statement in the paper body and (ii) the statistical evidence that supports or refutes that hypothesis. This formulation induces a challenging within-document retrieval setting in which many candidate paragraphs are topically related to the finding but differ in rhetorical role, creating hard negatives for retrieval and extraction. Using a two-stage retrieve-and-extract framework, we conduct a controlled study of retrieval design choices, varying context quantity, context quality (standard Retrieval Augmented Generation, reranking, and a fine-tuned retriever paired with reranking), as well as an oracle paragraph setting to separate retrieval failures from extraction limits across four Large Language Model extractors. We find that targeted context selection consistently improves hypothesis extraction relative to full-text prompting, with gains concentrated in configurations that optimize retrieval quality and context cleanliness. In contrast, statistical evidence extraction remains substantially harder. Even with oracle paragraphs, performance remains moderate, indicating persistent extractor limitations in handling hybrid numeric-textual statements rather than retrieval failures alone.

信息抽取科学文献大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。