用大模型自动提取文献数据,专家把关+模型互验提升准确性
Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models
- 让大模型自动生成提示词,效果接近人工专家设计
- 多模型交叉验证可减少错误,生成数据与专家判断高度一致
- 适合科研数据规模化整理,需保留人类审核环节
从科研论文中准确提取复杂、上下文相关的数据既耗时又费力。本文研究了前沿浏览器型大语言模型在提取高度上下文化信息方面的表现。通过四个递进流程发现:1)使用专家精心设计的提示词,多数先进LLM能有效提取数据,但对科学语境理解仍有不足;2)仅给简单指令,大模型即可自主生成提示词,效果几乎等同于专家撰写;3)自主发现文献困难,代理系统常遗漏或虚构参考文献;4)大模型可依据公开指南生成新数据集,与人工专家评分高度吻合,但仍需人类参与校验。这些结果定义了一种可审计的分工模式:专家设定证据标准,模型交叉验证重复提取,研究人员处理争议案例,为在不放弃专家控制的前提下实现科学数据整理规模化提供了可行路径。
原文摘要 · Abstract (English)
Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We demonstrate four escalating workflows, 1) given an expert curated prompt and research articles, most frontier LLMs perform well at data extraction, however can struggle with interpreting scientific context and nuance, 2) given simple instructions, LLMs can author their own prompts which were almost as eNective as expert-written prompts, 3) autonomous discovery of research literature was diNicult, agents either missed or hallucinated references, and 4) LLMs can create new datasets from published guidelines that closely match human-expert judges, but still require a human-in-the-loop. Together, these findings define an auditable division of labour in which experts specify the evidence standard, models cross-check repeated extractions and researchers resolve disputed cases, providing a practical route to scaling scientific data curation without relinquishing expert oversight.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。