arXiv:2603.23849cs.IR2026-03

用大模型从文献中提取病毒突变信息,解决科学数据构建难题

VILLA: Versatile Information Retrieval From Scientific Literature Using Large LAnguage Models

  • 设计多步检索增强生成框架VILLA,处理开放性突变提取任务
  • 构建629个流感病毒突变数据集,覆盖239篇文献,作为新基准
  • 在复杂任务上超越现有方法,适合病毒学与生物信息学研究者

高质量标注数据集的缺乏制约了人工智能在科研中的应用。利用大语言模型(LLM)从文献中进行科学信息抽取(SIE)正成为自动化构建数据集的有效途径。然而,现有基于LLM的SIE方法和评测研究多聚焦于生物医学、化学等宽泛领域,局限于选择题形式,且仅针对短篇、格式规范的文本。对复杂、开放性任务中SIE潜力的探索仍严重不足。本研究以此前几乎未被关注的病毒学为切入点,设计了一项独特的开放式SIE任务:从文献中提取能改变病毒与宿主相互作用的突变。我们提出名为VILLA的多步检索增强生成(RAG)框架,并构建了一个包含629个突变、来自10种甲型流感病毒蛋白、覆盖239篇科学文献的新数据集,作为该任务的基准。实验表明,VILLA在新型综合评估中显著优于基线RAG及当前最先进的RAG与代理类工具。

原文摘要 · Abstract (English)

The lack of high-quality ground truth datasets to train machine learning (ML) models impedes the potential of artificial intelligence (AI) for science research. Scientific information extraction (SIE) from the literature using LLMs is emerging as a powerful approach to automate the creation of these datasets. However, existing LLM-based approaches and benchmarking studies for SIE focus on broad topics such as biomedicine and chemistry, are limited to choice-based tasks, and focus on extracting information from short and well-formatted text. The potential of SIE methods in complex, open-ended tasks is considerably under-explored. In this study, we used a domain that has been virtually ignored in SIE, namely virology, to address these research gaps. We design a unique, open-ended SIE task of extracting mutations in a given virus that modify its interaction with the host. We develop a new, multi-step retrieval augmented generation (RAG) framework called VILLA for SIE. In parallel, we curate a novel dataset of 629 mutations in ten influenza A virus proteins obtained from 239 scientific publications to serve as ground truth for the mutation extraction task. Finally, we demonstrate VILLA's superior performance using a novel and comprehensive evaluation and comparison with vanilla RAG and other state-of-the art RAG- and agent-based tools for SIE.

信息抽取大模型应用病毒学RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。