arXiv:2606.08071cs.CL2026-06

构建首个大规模外科理解评估基准,测试大模型在手术场景下的推理能力。

SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models

论文配图:SurgiQ: A Large-Scale Multi-Domain Benchmark for Evaluating Surgical Understanding in Large Language Models
图 1 · 摘自论文原文
  • 基于教材与文献生成1.3万道四选一多题,覆盖六大学科领域。
  • 最佳模型准确率68.1%,多数小模型仅达随机水平(25%)。
  • 通用大模型如Qwen2.5表现优于专业医学模型,提示医疗专化不足。

当前大语言模型在外科领域的可靠评估仍不充分。通用医学基准侧重临床知识,而外科需处理流程推理、决策权衡、否定判断及多个合理术式选择。我们提出SurgiQ,一个纯文本、基于来源的基准,包含13,055道四选项多选题,覆盖六个外科领域和四种题型:病例类、推理类、最优选项类、否定类。数据源自外科教材、开放获取论文与考试材料,通过多阶段生成、验证与专家审核流程构建。我们在统一对数似然协议下评估了35个开源大模型。结果表明仍有显著提升空间:小型模型普遍接近25%随机基线,最佳模型达68.1%准确率。通用模型(尤其是Qwen2.5)优于多数生物医学模型,说明现有医学专业化尚未提供足够广的外科覆盖。校准与错误分析显示,即使强模型也会对临床上合理的干扰项做出自信误判,亟需更可靠、更全面的外科大模型评估体系。

原文摘要 · Abstract (English)

Reliable evaluation of large language models in surgery remains underdeveloped. Broad medical benchmarks test clinical knowledge, while surgery requires procedural reasoning, management trade-offs, negation handling, and selection among plausible operative decisions. We present SurgiQ, a text-only, source-grounded benchmark of 13,055 four-option multiple-choice questions spanning six surgical domains and four question formats: case-based, reasoning, best-option, and negative. SurgiQ is constructed from surgical textbooks, open-access papers, and examination material using a multi-stage generation, verification, and expert-audit pipeline. We evaluate 35 open-weight LLMs under a unified log-likelihood protocol. Our results show substantial remaining headroom: smaller models often remain near the 25\% random baseline, while the best model reaches 68.1\% accuracy. General-purpose models, especially Qwen2.5, outperform most biomedical models, suggesting that current medical specialization does not yet provide sufficiently broad surgical coverage. Calibration and error analysis further show that even strong models make confident mistakes on clinically plausible distractors, motivating more reliable and broader surgical LLM evaluation.

外科理解大模型评估多选题LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。