用大模型自动提取疾病传播模型论文信息,准确率超77%。
Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models

- 构建LLM流水线,从536篇论文中提取建模相关数据。
- GPT-5.0论文级准确率达81.67%,领域级最高达100%。
- 模型间一致性强可作质量判断,适合建模领域研究者参考。
大型语言模型(LLMs)的进展为简化甚至自动化研究流程(如系统性文献综述)提供了新机遇。本研究开发了一套基于LLM的管道,用于从536篇同行评审的基于代理的建模论文中提取与模型相关的信息,并与人工完成的系统性文献综述结果进行对比。结果显示,GPT-4.1在论文级准确率为约77.95%,GPT-5.0达81.67%。在领域级准确率上,范围为32.40%至100.00%,其中复杂或主观性较强的领域表现较弱。重要的是,我们发现模型间的一致性可作为输出质量的潜在指标:低一致性可能暗示幻觉,而高一致性但低准确率则可能指向人类数据集中的噪声或错误。总体而言,本研究为提示工程提供了实践洞见,揭示了在建模与仿真领域使用LLMs进行全规模系统性文献综述的潜力与局限。
原文摘要 · Abstract (English)
Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extracting model-relevant information from 536 peer-reviewed agent-based modeling papers. We compare the results with those of a human-conducted SLR. Our results show paper-level accuracies of approximately 77.95% for GPT-4.1 and 81.67% for GPT-5.0. Field-level accuracy ranges from 32.40% to 100.00%, with more complex or subjective fields performing less reliably. Importantly, we find that agreement between LLMs is a potential indicator of output quality: low agreement may signal hallucinations, whereas high agreement combined with low accuracy may point to noise or errors in the human dataset. Overall, our study provides practical insights into prompt development and highlights both the potential and limitations of using LLMs for full-scale SLRs in the modeling and simulation domain.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。