arXiv:2609.07601cs.CLcs.AI2026-09

测试大模型从真实病历中做产科决策的能力,发现证据提取是关键瓶颈。

ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making

论文配图:ObGynLongBench: Revealing the Evidence-to-EHR Gap in Longitudinal EHR Decision-Making
图 1 · 摘自论文原文
  • 构建产科长时序病历评测集,支持逐级信息获取评估
  • 模型在直接给证据时准确率高,从完整病历中提取证据时下降明显
  • 主动搜索策略表现最优,凸显个性化证据利用的挑战

大型语言模型在个性化医疗助手中的应用日益受到关注。然而,现有医学评测大多依赖静态问答和预选证据,无法判断模型能否从真实的纵向电子健康记录(EHR)中做出可靠临床决策。为填补这一空白,我们提出ObGynLongBench,一个基于规则的产科与妇科长时序EHR评测基准,包含来自976个真实妊娠病历的1,500个临床决策点案例及可追溯规则。每个案例关联患者、妊娠时间点与决策前信息边界,支持仅证据、单次就诊记录和全病史记录三种评估模式。评估17个LLM发现存在显著的Evidence-to-EHR差距:当证据直接提供时模型表现良好,但需从当日记录或完整决策前病历中提取证据时,准确率大幅下降。进一步分析表明,证据利用是主要瓶颈——随着病历上下文变长、证据需求复杂度提高,性能下降;早期失败常预示同一病历后续失败。主动搜索代理在所有访问策略中表现最佳,凸显个性化证据提取是实现可靠医疗助手的核心挑战。资源已开源:https://github.com/xiangjun2003/ObgynLongbench。

原文摘要 · Abstract (English)

The application of large language models (LLMs) to personalized medical assistants has garnered growing interest. However, existing medical benchmarks largely rely on static question answering with pre-selected evidence, leaving unclear whether LLMs can make reliable clinical decisions from real longitudinal electronic health records (EHRs). To bridge this gap, we introduce ObGynLongBench, a rule-grounded long-context EHR benchmark for obstetric and gynecologic decision-making, comprising 1,500 clinical decision-point cases from 976 real pregnancy EHR histories and traceable rules. Each case is anchored to a patient, a pregnancy-timeline point, and a pre-decision information boundary, enabling Evidence-only, Visit-level EHR, and History-level EHR evaluation. Evaluating 17 LLMs reveals a substantial Evidence-to-EHR Gap: models perform well when evidence is directly provided, but accuracy drops when evidence must be extracted from same-day records or full pre-decision EHR histories. Further analyses identify evidence utilization as a key bottleneck: performance decreases with longer EHR contexts and more complex evidence requirements, and earlier failures often predict later failures within the same patient history. Finally, active-search agents perform best among EHR access strategies, highlighting patient-specific evidence utilization as a central challenge for reliable personalized medical assistants. Resources are available at https://github.com/xiangjun2003/ObgynLongbench.

医疗AI长时序病历大模型评测证据提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。