用反事实测试发现临床大模型隐藏的响应能力差异
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents

- 设计因果敏感度评分,通过五类临床变量扰动测试模型响应变化
- 六款前沿模型在覆盖率和响应性上排名完全相反,最低覆盖模型反而最敏感
- 揭示所有模型在手术状态变化时均失效,暴露覆盖率指标无法发现的安全盲区
两个临床AI系统在覆盖率评分上几乎相同,但面对患者输入变化时行为迥异:一个能根据新临床信号调整建议,另一个则保持输出不变。我们提出因果敏感度评分(CSS),一种预注册的干预型评估指标,对224例肿瘤会诊病例沿五个临床有意义维度(生物标志物翻转、既往治疗失败、生物标志物移除、手术状态变更、分期扰动)进行干预,并以{0, 0.5, 1.0}尺度评分模型是否按预设正确方向更新推荐。与基于覆盖率的共识匹配分数(CMS)对比,六款来自三所机构的前沿模型在单次推理下排名近乎完全相反:所有模型排名变动,CMS最差者成为CSS最佳,一名中上位CMS模型在CSS中垫底。进一步发现普遍安全盲点:所有模型在手术状态干预下表现极差(家族D最多仅17.2% CSS),而CMS未揭示此问题。该指标亦适用于工具使用代理:在ReAct式实验中,工具使用使五款模型的CSS提升2.5至20.3个百分点,但最低分模型仍检索相同病历片段且未更新建议,暴露出结构性响应缺陷——唯有反事实评估才能揭示。跨评审员复现与三名医学专业人员验证确认整体结论。干预型预注册指标如CSS可补充覆盖率评估,捕捉覆盖率忽略的响应能力,为未来代理式强化学习提供潜在密集奖励信号。
原文摘要 · Abstract (English)
Two clinical AI systems can score nearly identically on coverage-based rubrics yet behave radically differently when their patient inputs change: one updates its recommendations to match the new clinical signal, while the other produces the same output regardless. We introduce the Causal Sensitivity Score (CSS), a pre-registered interventional metric that mutates oncology tumor-board cases along five clinically meaningful dimensions - biomarker flips, prior-treatment failures, biomarker removals, surgery-status changes, and stage perturbations - and scores whether each model updates its recommendations in the pre-registered correct direction using a {0, 0.5, 1.0} scale. Benchmarked against the Consensus Match Score (CMS), a coverage-based weighted recall metric, six frontier models from three labs evaluated in single-shot inference across 224 cases rank in nearly opposite orders: all six models change rank, the CMS-worst model becomes CSS-best, and one upper-mid CMS model ranks last on CSS. We further surface a universal safety blind spot: every frontier model fails on surgery-status interventions (at most 17.2% CSS on Family D), a finding CMS does not expose. The metric also transfers to tool-using agents: in a ReAct-style experiment, tool use improves CSS for five of six models (+2.5 to +20.3 percentage points), yet the lowest-CSS model retrieves the same chart sections and still fails to update its recommendations - revealing a structural responsiveness deficit visible only under counterfactual evaluation. Cross-judge replication and three-rater medical-professional validation confirm the aggregate findings. Interventional pre-registered metrics like CSS complement coverage-based evaluation for clinical AI agents: they capture responsiveness that coverage metrics miss and offer a candidate dense reward signal for future agentic RL systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。