arXiv:2605.22612cs.CYcs.AI2026-05

医疗大模型评估需揭示隐含假设,否则部署表现难预测。

Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions

  • 区分任务与结果两类用户假设,任务可从对话数据验证
  • 实证发现评估与部署差距中,任务与结果各占约一半
  • 提出BenchmarkCards和分阶段评估,系统检验假设

评估基准对医疗AI至关重要,但不足以预测实际部署表现。我们指出,评估-部署差距的根源不在于基准设计不佳,而在于无法通过基准暴露的用户交互隐含假设。为此,我们提出将假设分为两类:仅需对话数据即可检验的‘任务’假设,以及需结果数据与行为研究才能验证的‘结果’假设。后者依赖人类行为,是基准难以直接观测的。通过回顾一项医疗随机对照试验作为案例,我们发现该差距自然分解为大致相当的任务与结果部分。为此,我们提出两项贡献:一是设计BenchmarkCards以记录假设,二是提出分阶段评估流程,系统性测试假设并评估性能。

原文摘要 · Abstract (English)

Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises not because of poorly designed benchmarks, but from implicit assumptions about how users interact with models that cannot be surfaced from benchmarks alone. To make this precise, we propose a classification of assumptions into two categories: task, which can be tested from conversation data alone, and outcome, which requires outcome data and behavioral studies for testing. Critically, outcome assumptions depend on human behavior, something that even well-designed benchmarks cannot directly observe. To demonstrate the operationality of this framework, we retrospectively analyze a healthcare RCT as a case study and find that the gap naturally separates into task and outcome gaps of roughly equal size. To address this, we make two contributions: first, we propose BenchmarkCards, an artifact that documents assumptions, and second, we propose staged evaluation, a procedure that systematically tests assumptions and evaluates performance.

医疗AI评估框架用户假设

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。