arXiv:2606.02568cs.AIcs.CL2026-06被引 2

构建可交互的长期住院模拟环境,评估大模型临床决策能力

ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents

论文配图:ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents
图 1 · 摘自论文原文
  • 设计多阶段交互式医疗环境,模拟真实住院诊疗流程
  • 7个模型最高决策F1仅0.31,管理决策准确率远低于诊断(0.17 vs 0.51)
  • 首次量化信息获取效率与决策质量的差距,适合评估临床AI代理

临床实践并非在有限选项中做选择:医生需逐步收集异构信息,在不确定性下做出连续且不可逆的决策。现有静态基准无法充分检验,而现有交互式医疗基准至少在某一方面有所妥协。我们提出ClinEnv,一个以纵向住院模拟为范式的交互式基准,评估大语言模型作为住院医师在真实入院病例上的表现。每个病例自动构建为有序的决策阶段序列;每阶段模型需主动向四个专业代理查询后,再决定用药、操作和诊断。ClinEnv通过确定性本体匹配评估决策结果,并衡量信息获取过程。在七种模型中,最强模型的决策F1仅为0.31,且结果质量与过程质量显著脱钩。难点集中在管理决策和后期阶段,模型对出院诊断的恢复较可靠(F1=0.51),但对管理动作的准确性极低(F1=0.17),且随病例推进持续发出冗余查询。ClinEnv使以往仅凭结果评估难以察觉的信息获取差距得以直接测量。

原文摘要 · Abstract (English)

Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to sequential, irreversible decisions under uncertainty. Static benchmarks cannot probe and existing interactive medical benchmarks each compromise on at least one of them. We present ClinEnv, an interactive benchmark that evaluates LLMs as attending physicians over real inpatient admissions under a paradigm we term Longitudinal Inpatient Simulation. Each case is automatically constructed into an ordered sequence of decision stages; at every stage the model must actively query four specialized agents before committing to medications, procedures, and diagnoses. ClinEnv scores both what the model decides, through deterministic ontology-grounded matching, and how it gathers information. Across seven models, the strongest reaches only 0.31 decision F1, and outcome quality is sharply decoupled from process quality. Difficulty concentrates in management decisions and later stages, where models recover discharge diagnoses far more reliably than management actions (0.51 vs. 0.17 F1) and continue to issue redundant queries as cases progress. ClinEnv makes this information-acquisition gap, invisible to outcome-only evaluation, directly measurable.

临床决策大模型评估医疗模拟交互式基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。