arXiv:2608.13786cs.IRcs.AI2026-08

AI聊天机器人找医学研究,效果受模型和用户角色影响,更偏爱大样本研究。

Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions

  • 用三款主流AI模型模拟患者、医生等角色提问,评估其检索临床研究的能力。
  • 平均仅召回39.2%的权威研究,但误引了5.0%的被排除研究,大样本试验更易被选中。
  • 适合关注AI在医学信息检索中局限性的研究人员或临床决策支持开发者。

大型语言模型(LLM)聊天机器人正被用于回答临床问题并引用相关研究。现有研究多关注引用伪造,却忽视对检索出研究质量及其选择因素的评估。本研究评估了Claude Sonnet 5、Gemini 3.1 Pro和ChatGPT GPT-5.5三款通用模型。基于2026年Cochrane系统综述第6、7期中的20个问题,模拟患者、临床医生及循证研究者角色,每模型在每角色下独立测试4次,共生成720条回复。要求模型提供原始临床研究引用,并与Cochrane综述的纳入/排除研究集对比。结果显示,模型平均召回39.2%±29.8%的纳入研究,同时引用5.0%±9.4%的排除研究。召回率显著受模型与用户角色影响:ChatGPT表现最佳(63.1%±29.5%),优于Claude(37.0%±23.8%)和Gemini(17.3%±13.1%;p=2.0×10⁻⁵);研究者角色召回率高于医生或患者角色(42.8%±30.8% vs. 38.6%±28.9% vs. 36.1%±29.3%;p=2.0×10⁻⁵)。控制发表年份、年均引用量及开放获取状态后,样本量是唯一独立显著预测因子(每单位对数样本量增加,优势比1.80,95% CI 1.37–2.36,p=2.34×10⁻⁵)。结果表明,尽管模型可捕捉部分专家识别的研究,但性能差异明显,且存在对大样本研究的偏好。

原文摘要 · Abstract (English)

Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\pm$ 29.5% vs. 37.0% $\pm$ 23.8% vs. 17.3% $\pm$ 13.1%; $p=2.0\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\pm$ 30.8% vs. 38.6% $\pm$ 28.9% vs. 36.1% $\pm$ 29.3%; $p=2.0\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.

医学问答AI检索大模型评估样本量偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。