用大模型自动提取医学文献数据,准确率超80%
OpenExtract: Automated Data Extraction for Systematic Reviews in Health
- 调用大模型分析论文段落,预测数据条目
- 在数字健康综述中实现精度与召回率均超0.8
- 开源工具适合科研人员加速文献综述
本研究提出OpenExtract,一个用于大规模系统性文献综述的开源自动化数据提取流程。该流程通过调用大语言模型(LLMs),基于科学论文的相关段落预测数据条目。为验证OpenExtract的有效性,我们将其应用于数字健康领域的系统性文献综述,并与人工研究人员的结果进行对比。结果显示,OpenExtract在该任务中的精度和召回率均超过0.8,表明其能高效且准确地实现数据自动提取。项目地址:https://github.com/JimAchterbergLUMC/OpenExtract。
原文摘要 · Abstract (English)
This study presents OpenExtract, an open-source pipeline for automated data extraction in large-scale systematic literature reviews. The pipeline queries large language models (LLMs) to predict data entries based on relevant sections of scientific articles. To test the efficacy of OpenExtract, we apply it to a systematic literature review in digital health and compare its outputs with those of human researchers. OpenExtract achieves precision and recall scores of > 0.8 in this task, indicating that it can be effective at extracting data automatically and efficiently. OpenExtract: https://github.com/JimAchterbergLUMC/OpenExtract.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。