评测大模型在医学文献数据提取中的表现,给出自动化程度的实用分级建议。
What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction
- 用定制提示词提升关键信息召回率最多达15%
- 三类模型均高精度但低召回,漏提重要数据
- 提出三档自动化指南,适配不同复杂度任务
自动化从随机对照试验全文中提取元分析所需数据仍面临重大挑战。本研究评估了三种大语言模型(Gemini-2.0-flash、Grok-3、GPT-4o-mini)在高血压、糖尿病和骨科三个医学领域中,对统计结果、偏倚风险评估及研究特征等任务的表现。采用四种提示策略(基础提示、自我反思提示、模型集成与定制提示)以提升提取质量。所有模型均表现出高精确率,但普遍存在召回率低的问题,常遗漏关键信息。研究发现,定制提示最有效,可将召回率提升最高达15%。基于此,我们提出一套三层次的使用指南,根据任务复杂性和风险等级,匹配相应的自动化水平。本研究为真实世界元分析中的自动化数据提取提供了实践指导,实现大模型效率与专家审核之间的平衡。
原文摘要 · Abstract (English)
Automating data extraction from full-text randomised controlled trials (RCTs) for meta-analysis remains a significant challenge. This study evaluates the practical performance of three LLMs (Gemini-2.0-flash, Grok-3, GPT-4o-mini) across tasks involving statistical results, risk-of-bias assessments, and study-level characteristics in three medical domains: hypertension, diabetes, and orthopaedics. We tested four distinct prompting strategies (basic prompting, self-reflective prompting, model ensemble, and customised prompts) to determine how to improve extraction quality. All models demonstrate high precision but consistently suffer from poor recall by omitting key information. We found that customised prompts were the most effective, boosting recall by up to 15\%. Based on this analysis, we propose a three-tiered set of guidelines for using LLMs in data extraction, matching data types to appropriate levels of automation based on task complexity and risk. Our study offers practical advice for automating data extraction in real-world meta-analyses, balancing LLM efficiency with expert oversight through targeted, task-specific automation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。