arXiv:2605.13542cs.AIcs.CL2026-05

新基准RealICU让大模型在真实重症数据中检验推理能力,超越简单模仿医生行为。

RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation

论文配图:RealICU: Do LLM Agents Understand Long-Context ICU Data? A Benchmark Beyond Behavior Imitation
图 1 · 摘自论文原文
  • 用事后回顾标注的完整病程数据构建评估标准,更贴近临床真实决策场景。
  • 94名患者共930个30分钟片段标注,11862个扩展数据由医生验证的AI生成。
  • 发现模型存在安全与召回权衡、早期判断锚定偏差等关键缺陷,适合重症AI研究者参考。

重症监护室(ICU)产生长序列、高密度且动态演化的临床数据,医生需在时间压力下反复评估患者状态,凸显可靠AI辅助决策的迫切需求。现有ICU评测通常将历史医生行为视为真值,但这些行为基于不完整信息和有限时序上下文,可能非最优,难以真实评估AI的推理能力。我们提出RealICU,一个基于事后回溯标注的基准,用于评估大语言模型(LLMs)在真实ICU条件下的表现,其中标签由资深医生在完整患者轨迹回顾后生成。我们设计四个医生驱动任务:评估患者状态、急性问题、推荐措施及可能引发不安全后果的警示动作。每个病程按30分钟窗口划分,发布两个数据集:RealICU-Gold包含94名MIMIC-IV患者共930个窗口的标注;RealICU-Scale通过一位经医生验证的LLM回溯标注器(Oracle)扩展至11,862个窗口。现有大模型,包括带记忆增强的模型,在RealICU上表现不佳,暴露出两类失败模式:临床建议中的召回-安全权衡,以及对早期患者判断的锚定偏差。我们进一步引入ICU-Evo以研究结构化记忆代理,虽提升长时推理能力,但仍未能完全消除安全问题。RealICU为衡量和改进高风险医疗场景中的AI序列决策支持提供了临床根基的测试平台。

原文摘要 · Abstract (English)

Intensive care units (ICU) generate long, dense and evolving streams of clinical information, where physicians must repeatedly reassess patient states under time pressure, underscoring a clear need for reliable AI decision support. Existing ICU benchmarks typically treat historical clinician actions as ground truth. However, these actions are made under incomplete information and limited temporal context of the underlying patient state, and may therefore be suboptimal, making it difficult to assess the true reasoning capabilities of AI systems. We introduce RealICU, a hindsight-annotated benchmark for evaluating large language models (LLMs) under realistic ICU conditions, where labels are created after senior physicians review the full patient trajectory. We formulate four physician-motivated tasks: assess Patient Status, Acute Problems, Recommended Actions, and Red Flag actions that risk unsafe outcomes. We partition each trajectory with 30-min windows and release two datasets: RealICU-Gold with 930-window annotations from 94 MIMIC-IV patients, and RealICU-Scale with 11,862 windows extended by Oracle, a physician-validated LLM hindsight labeler. Existing LLMs including memory-augmented ones performed poorly on RealICU, exposing two failure modes: a recall-safety tradeoff for clinical recommendations, and an anchoring bias to early interpretations of the patient. We further introduce ICU-Evo to study structured-memory agents that improves long-horizon reasoning but does not fully eliminate safety failures. Together, RealICU provides a clinically grounded testbed for measuring and improving AI sequential decision-support in high-stakes care. Project page: https://chengzhi-leo.github.io/RealICU-Bench/

重症医疗大模型评估长期推理临床决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。