arXiv:2608.28592cs.AI2026-08

九款前沿大模型在癌症决策中普遍无法做出正确路径选择,暴露了系统性盲区。

A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making

论文配图:A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making
图 1 · 摘自论文原文
  • 构建癌症决策基准测试,评估模型在指南路径中的判断能力。
  • 42.1%的决策点无一模型答对,关键错误在于路径选择而非推理细节。
  • 模型越果断越易误判,需设计能识别能力边界并转交医生的架构。

大型语言模型(LLMs)在医学知识考试中表现优异,但真实癌症诊疗并非知识测验,而是基于指南、逐步决策与不确定性下的判断。现有评测多聚焦事实记忆,未揭示前沿模型是否共享决策路径盲点,且合并模型也难以弥补。我们构建了癌症决策边界基准(ODBB),涵盖NCCN指南和结直肠癌病例中的2,005个决策点,并评估了2025年6月至2026年4月间发布的九款前沿大模型(四款闭源,五款开源)。采用完全确定性评分器(零次大模型推理)将输出分为14类失败,由两名肿瘤科医生独立验证(加权科恩卡帕系数分别为0.939和0.790,样本量225项)。将九款模型视为集成超模型,42.1%(置信区间40.0–44.3%)的全部项目——其中NCCN项目占35.7%(1,586项),结直肠癌案例占66.4%(419项)——均无模型正确回答。失败集中在指南路径选择阶段之前,而非其内部推理,表明临床元判断存在系统性盲区,可能需架构层面干预而非更多数据训练。两款强调决断力的模型(GPT-5.5、Gemini 3.1 Pro Preview)比七款谨慎模型更频繁做出不安全承诺,频率高出三至五倍,且未提升准确率。在3–9%的案例中,模型虽说出正确下一步,却未明确执行——体现的是决策失败,而非知识不足。模型质量已不再是临床部署的主要瓶颈,真正约束是‘单模型可独立决策’的假设。进步需依赖能识别自身能力边界并自动转交医生的架构。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)--2,005 oncology decision points across NCCN guidelines and colorectal cancer cases--and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 2026. A fully deterministic scorer (zero LLM inference) classified outputs into 14 failure types, independently validated by two oncologists (Cohen's weighted $κ$ = 0.939 and 0.790) on a 225-item stratified sample. Treating the nine as a pooled super-model, 42.1% (Wilson 95% CI 40.0--44.3%) of all items--35.7% of the 1,586 NCCN items and 66.4% of the 419 colorectal-cancer cases--were answered correctly by none, with failures concentrated in choosing between guideline pathways before reasoning within any: a consistent blind spot in clinical meta-judgment that likely requires architectural intervention rather than more training data. Two models tuned for decisiveness (GPT-5.5, Gemini 3.1 Pro Preview) made unsafe commitments three to five times more often than the seven cautious models without scoring higher. In 3--9% of items, models stated the correct next clinical step yet did not commit to it--failures of decision, not knowledge. Model quality is no longer the primary bottleneck for clinical LLM deployment; the binding constraint is the assumption that any single model can be the sole basis for a clinical decision. Progress requires architectures that detect when a model reaches its competence boundary and route the decision to a clinician.

癌症决策大模型边界临床应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。