大模型可精准提取微生物致癌证据,媲美专家水平。
Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications

- 用结构化问卷评估大模型在24篇论文中提取和评估证据的能力。
- GPT-5和GPT-5 Nano与专家一致性接近,表现无差异。
- 方法学评估和文本矛盾识别仍是模型薄弱环节。
已确认的致癌微生物对癌症负担贡献显著。发现新型微生物致癌性可能带来降低疾病负担的新策略。然而,相关证据分散,人类难以全面整合。大语言模型(LLM)有望实现可扩展的、专家级的系统性证据合成,以识别微生物-癌症关联;但该能力尚未被证实。研究招募领域专家,构建数据集以基准测试LLM(Gemini 2.5 Pro、Gemini 2.5 Flash、GPT-5、GPT-5 Nano)在24篇文献中的表现,以MMTV-LV与乳腺癌为案例。设计包含MCQ、Likert量表、多选和自由文本共77项的结构化模板进行证据提取与评价。通过新指标评估专家与各模型间每题的一致性。通过比较专家间及专家-模型一致性的分布,判断模型是否表现出如额外专家般提升或维持一致性。自由文本响应亦进行定性评估。所有题型中,模型响应与专家高度一致,其中GPT-5与GPT-5 Nano的表现分布与专家无法区分。Gemini模型表现相似,但在应用微生物致癌标准上明显更宽松。幻觉罕见。方法学评价和全文内部矛盾识别是模型最持续的弱点。结论:GPT-5与GPT-5 Nano在结构化领域文献评估任务中与专家无差异,支持其用于自动化系统性证据合成。但方法学评价与全文矛盾识别仍需改进。
原文摘要 · Abstract (English)
Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensively synthesize. LLMs may enable scalable, expert-level systematic evidence synthesis to identify microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Domain experts were recruited to create a dataset to benchmark LLM performance (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano) on 24 research papers using MMTV-LV and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal, consisting of MCQ, Likert-scale, multi-select, and free-text question types (77 items across 24 papers). Agreement between (1) experts and (2) experts and each LLM was determined per question instance using novel metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement distributions to determine whether LLMs behaved as additional experts by increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Across all question types, LLM responses aligned closely with experts, with GPT-5 and GPT-5 Nano achieving score distributions indistinguishable from experts. Gemini models behaved similarly but were significantly more lenient in applying microbial oncogenesis criteria. Hallucinations were rare. Methodological appraisal and identification of contradictions within full-texts were the most persistent LLM vulnerabilities. GPT-5 and GPT-5 Nano were indistinguishable from experts on structured domain research paper evaluation tasks. This supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-texts remain weaknesses requiring strengthening.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。