arXiv:2608.07796cs.AI2026-08

评测医疗大模型在真实病历中的推理能力,强调证据可信与合理拒答。

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR

论文配图:CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR
图 1 · 摘自论文原文
  • 构建基于真实病历的临床推理评测框架,模拟完整诊疗调查过程。
  • 16个智能体四类判断准确率65.3%-76.1%,但缺陷免罚准确率低4.8-14.8点。
  • 首次统一评估证据溯源、政策遵循、过程合规与合理拒答,适合临床部署研究者。

大型语言模型在医学知识评测中表现优异,但可靠临床应用需其能在异构、纵向电子病历中开展可辩护的调查:明确所需证据、检索并整合结构化与自由文本数据、结论基于可验证证据,并对无法可靠解决的案例合理拒答。本文提出CliniCARE-Bench(电子病历中临床推理的校准审计),一个回顾性临床审计基准,包含25个由临床医生验证的场景,生成750例基于真实患者数据的MIMIC-IV病例。系统在受控日志工具环境中进行记录检索、计算和政策访问,返回四种结论——是、否、因数据缺失而不确定、因医学模糊而不确定,后两者区分证据缺失与医学争议。除结论准确性外,还评估患者证据关联性、政策依据、流程合规性、校准拒答、可靠性与效率,均以独立多模型仲裁及临床委员会校准的参考结论为基准。所有检索、计算与报告均可回放,调查轨迹可审查、可评分。据我们所知,CliniCARE-Bench是首个面向部署的临床智能体基准,统一评估真实纵向病历调查、声明级证据溯源、政策使用、流程合规与校准拒答。16个代理系统四类准确率在65.3%至76.1%之间,但原始准确率高估了调查质量;无缺陷准确率(仅当结论正确且无禁止捷径时计分)低4.8至14.8个百分点,且重新排序了排行榜。

原文摘要 · Abstract (English)

Large language models perform strongly on medical knowledge benchmarks, but reliable clinical deployment requires agents to conduct defensible investigations over heterogeneous, longitudinal records: determining what evidence is needed, retrieving and reconciling structured and free-text data, grounding conclusions in verifiable evidence, and deferring cases that cannot be resolved reliably. We introduce CliniCARE-Bench (Clinical Calibrated Audit of Medical Reasoning in EHR), a benchmark for retrospective clinical audit: 25 clinician-validated scenarios instantiated as 750 patient-specific cases over real-patient-derived MIMIC-IV data. Systems investigate each case through a governed, logged tool environment for record retrieval, computation, and policy access, and return one of four verdicts---Yes, No, Indeterminate: Lack of Data, or Indeterminate: Medically Ambiguous---the last two separating missing evidence from residual medical ambiguity. Beyond verdict accuracy, we score patient-evidence and policy grounding, process adherence, calibrated abstention, reliability, and efficiency against case-level reference verdicts produced by independent multi-model adjudication and calibrated against Clinical Board review. Every retrieval, computation, and report is replayable, so the investigation trace is inspectable and scorable. To our knowledge, CliniCARE-Bench is the first deployment-oriented clinical-agent benchmark to jointly evaluate real longitudinal EHR investigation, claim-level evidence grounding, governing-policy use, process adherence, and calibrated abstention within a common patient-level adjudication framework. Across 16 agentic systems, four-way accuracy spans 65.3-76.1%, but raw accuracy overstates investigation quality. Defect-free accuracy, which credits a verdict only when correct and free of prohibited shortcuts, is 4.8-14.8 points lower and reorders the leaderboard.

医疗AI病历推理可解释性评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。