真实临床提问测试显示,专业工具优于通用AI,专家评估更可信。
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

- 用620个真实医生提问做盲评,对比三大模型与专业工具。
- 专业工具在准确性和临床实用性上胜出25至39个百分点。
- 研究呼吁用真实场景和专科专家评估AI,适合医疗AI研发者参考。
医生每周向AI工具提出数百万个临床问题,但现有评估多基于假设或考试式问题,而非实际使用场景。本研究基于来自30个专科的620个真实临床点对点查询(Real-POCQi)及HealthBench中的187个问题,由来自36个州的149名执业医生开展盲法头对头比较,评估Claude Opus 4.8、Gemini 3.1 Pro、GPT-5.5三款前沿通用模型与专业临床工具OpenEvidence(OE)的表现。评估维度包括准确性、临床实用性、来源质量、可验证性与完整性。结果显示,专业工具在所有维度均得分最高;在主要分析中,胜率差距为25至39个百分点(p<0.001)。敏感性分析按引用展示、回答长度、OE用户状态及数据集来源分层后结果一致。同时发现,大模型裁判与专家裁判存在系统性差异,但总体认可最佳模型。结论指出:(i)AI评估应反映真实查询分布,且需匹配专科背景的专家评委;(ii)专业工具的优势并非通用模型不可替代,而是定向优化可带来显著性能提升。研究公开发布Real-POCQi基准及预设统计分析代码供复现。
原文摘要 · Abstract (English)
Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in practice. We report a blinded evaluation built on 620 Real-world Point-Of-Care Queries (Real-POCQi) submitted to the OpenEvidence (OE) platform by physicians spanning 30 specialties, as well as 187 questions from HealthBench. 149 practicing physicians across 36 states made head-to-head comparisons between answers from three frontier general-purpose models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5) and a specialized clinical tool (OE), with graders matched to each question's specialty. When comparing answers along five dimensions relevant to clinical decision support -- accuracy, clinical utility, source quality, verifiability, & completeness -- physicians scored the specialized tool highest on all axes; in the primary analysis on Real-POCQi, win differences (margins between win and loss rates) ranged from 25 to 39 percentage points (p<0.001). Results remained consistent in sensitivity analyses stratifying by citation display, answer length, OE-user status, and Real-POCQi versus HealthBench. In parallel, LLM judges were found to systematically differ from expert judges, though both generally agreed on the best model. These findings underscore two conclusions: (i) AI tool evaluations should reflect real-world query distributions and use expert judges that mirror the specialization defining modern medicine and (ii) the consistent advantage of the specialized tool over general-purpose models does not necessarily mean that the latter cannot serve similar purposes, but that targeted engineering and customization can yield meaningful gains in performance for its users. We release Real-POCQi as a public benchmark, as well as the prespecified statistical analysis for reproducing results of this study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。