arXiv:2607.10849cs.CL2026-07

Claude Fable 5在生物医学任务中拒绝率高达99.4%,但愿回答时准确率超越所有模型。

Capabilities of Claude Fable 5 on Biomedical Challenge Problems

论文配图:Capabilities of Claude Fable 5 on Biomedical Challenge Problems
图 1 · 摘自论文原文
  • 用固定答案键进行确定性评分,避免语言模型自评偏差。
  • 拒绝率在8.0%至99.4%间波动,排除拒绝项后准确率全面领先。
  • 拒绝行为分两类:基础科学内容和罕见病类型,体现安全策略差异。

前沿大模型在生物医学基准上评估日益普遍,但现有评估存在两大问题:旧基准接近饱和,且开放回答由其他语言模型评分。本文对Anthropic最新公开模型Claude Fable 5在八个生物医学基准(四个文本、四个多模态)上进行评估,采用固定答案键的确定性评分。对比包含两个Claude前代模型及GPT-5。所有结果表均单独记录拒绝情况。核心发现为:Fable 5在不同基准上拒绝率介于8.0%至99.4%之间,该现象在前代模型及GPT-5中均未出现。剔除拒绝项后,其准确率在所有任务中均超过或持平其他模型。识别出两种可区分的拒绝模式:一是集中在MedQA与MedXpertQA MM中的基础科学与机制类内容,通过两基准自身类别标签独立验证;二是罕见病基准(RareBench)中,先天代谢病表现几乎全被拒绝,而成人自身免疫病则基本不拒。因此,Fable 5在生物医学应用中的主要限制是响应意愿,而非能力本身。

原文摘要 · Abstract (English)

Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models. We evaluate Claude Fable 5, Anthropic's most capable publicly available model, across eight biomedical benchmarks, four text and four multimodal, using deterministic scoring against fixed answer keys throughout. We include two Claude predecessors and GPT-5 as baselines. Refusal is tracked as a distinct outcome in every result table. That decision produces the paper's central finding. Fable 5 refuses between 8.0% and 99.4% of questions depending on the benchmark, a pattern absent in both predecessors and in GPT-5. Once refused items are excluded from the denominator, Fable 5's accuracy exceeds or meets every other model on every benchmark in this study. We identify two distinguishable refusal patterns: one concentrating in basic-science and mechanism content across MedQA and MedXpertQA MM, confirmed independently on two benchmarks using each benchmark's own category labels; and a separate disease-domain pattern on RareBench, where inborn metabolic disease presentations are refused near-universally while adult-onset autoimmune presentations are not. The primary constraint on Fable 5's biomedical usefulness is willingness to engage, not capability once it does.

大模型评估生物医学拒绝行为安全策略

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。