arXiv:2608.15909cs.IRcs.CL2026-08

用大模型自动从文献中挖掘遗传研究所需队列,提升发现效率。

Large language model-assisted discovery of cohorts from scientific literature

论文配图:Large language model-assisted discovery of cohorts from scientific literature
图 1 · 摘自论文原文
  • 基于研究问题生成检索词,通过API自动抓取文献并用LLM提取队列名称。
  • 从5400个查询中找出188个候选队列,经人工筛选保留44个有效队列。
  • 可跨人群、表型和数据类型适配,弥补现有队列库的遗漏。

开展多研究分析需识别具有相关参与者、表型和数据模态的队列。传统方法依赖先验知识、队列目录和手动文献检索。本文开发了一种以问题为导向的互补框架,可自动搜索科学文献并提取明确的队列名称。该框架首先从可配置词汇和模板生成多个PubMed查询,通过PubMed API自动获取文献;随后利用大语言模型对标题和摘要进行筛查,采用针对研究问题定制的提示词提取队列名称,并经人工审阅去重。代码、提示词及示例输出已公开。作为应用案例,将该框架用于青少年攻击性遗传学研究:从5,400个生成查询中获取5,254条唯一记录,识别出188个候选队列;经预设标准(如年龄范围、基因数据可用性)人工筛选后保留44个符合条件的队列。基于LLM的自动化命名提取与人工标注者一致性在合理范围内。同时,使用相同问题搜索四个已知队列数据库,其合并结果仅包含44个合格队列中的27个,另有17个未被任何数据库覆盖。结论表明,该框架可将特定研究问题的词汇转化为可筛查的队列清单,适用于不同人群、表型、数据模态和研究设计,为结构化队列目录提供文献基础的补充。

原文摘要 · Abstract (English)

Background: Planning multi-study analyses requires identifying cohorts with the relevant participants, phenotypes, and data modalities. This process commonly relies on prior knowledge, cohort catalogues, and manual literature searches. We developed a complementary question-driven framework that searches relevant scientific literature and extracts explicit cohort names. Methods: The framework first generates multiple PubMed queries from configurable vocabularies and templates and retrieves the resulting scientific literature automatically through the PubMed API. A large language model then screens the retrieved titles and abstracts and extracts explicit cohort names using a prompt tailored to the research question. The extracted names are deduplicated with human review. Configurable code, prompts, and example outputs are available at https://gitlab.rz.uni-frankfurt.de/cap_molgenlab/literature-cohort-discovery. Evaluation: As a use case, we applied the framework to youth aggression genetics. From 5,400 generated PubMed queries, the framework retrieved 5,254 unique records and identified 188 candidate cohorts. Manual screening using predefined criteria, including participant age and genetic-data availability, retained 44 eligible cohorts. Automated LLM-based name extraction was within the agreement range of human annotators. We also searched four established cohort catalogues using the same research question. Their combined results contained 27 of the 44 eligible cohorts, while 17 were not returned by any cohort catalogue search. Conclusion: The framework converts research-question-specific vocabulary into screenable cohort inventories via a large, automated literature search. It can be adapted across populations, phenotypes, data modalities, and study designs, and provides a literature-based complement to curated cohort catalogues.

文献挖掘队列发现大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。