用大模型加速文献综述数据提取,但复杂信息准确率低。
Expediting data extraction using a large language model (LLM) and scoping review protocol: a methodological study within a complex scoping review
- 用评审协议引导大模型提取数据,提升效率。
- 简单信息提取准确率达83.3%~100%,复杂主观信息仅9.6%~15.8%。
- 适合需快速初筛的综述研究者,但结果需人工复核。
综述中的数据提取耗时耗力,本研究测试了使用Claude 3.5 Sonnet大模型结合评审协议,在一项复杂综述案例中从10个证据源提取数据的两种方法。结果显示,提取简单、明确的引用信息时准确率高达83.3%和100%;而提取复杂、主观的数据项时准确率仅为9.6%和15.8%。整体来看,两种方法精度均超过90%,但召回率低于25%,F1分数低于40%。性能受限于综述复杂性、开放式回答类型及方法设计。大模型反馈认为原始提取基本准确,仅建议微调部分条目(15条中4条引用信息,38条关键发现中8条)。但在含故意错误的数据集上,仅检测出39个错误中的2个(5%)。研究建议在类似应用中评估并报告大模型性能,并利用其反馈优化评审协议。
原文摘要 · Abstract (English)
The data extraction stages of reviews are resource-intensive, and researchers may seek to expediate data extraction using online (large language models) LLMs and review protocols. Claude 3.5 Sonnet was used to trial two approaches that used a review protocol to prompt data extraction from 10 evidence sources included in a case study scoping review. A protocol-based approach was also used to review extracted data. Limited performance evaluation was undertaken which found high accuracy for the two extraction approaches (83.3% and 100%) when extracting simple, well-defined citation details; accuracy was lower (9.6% and 15.8%) when extracting more complex, subjective data items. Considering all data items, both approaches had precision >90% but low recall (<25%) and F1 scores (<40%). The context of a complex scoping review, open response types and methodological approach likely impacted performance due to missed and misattributed data. LLM feedback considered the baseline extraction accurate and suggested minor amendments: four of 15 (26.7%) to citation details and 8 of 38 (21.1%) to key findings data items were considered to potentially add value. However, when repeating the process with a dataset featuring deliberate errors, only 2 of 39 (5%) errors were detected. Review-protocol-based methods used for expediency require more robust performance evaluation across a range of LLMs and review contexts with comparison to conventional prompt engineering approaches. We recommend researchers evaluate and report LLM performance if using them similarly to conduct data extraction or review extracted data. LLM feedback contributed to protocol adaptation and may assist future review protocol drafting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。