arXiv:2608.17075cs.CLcs.AI2026-08

用多模型协作预测未来诊断编码,提升临床决策支持可信度。

Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting

论文配图:Foundation Agents Meet Agentic Deep Research: Evidence-Grounded Clinical Code Forecasting
图 1 · 摘自论文原文
  • 融合EHR基础模型与语言模型,分步生成候选诊断代码
  • 在MIMIC-III/IV上达到24.60%/35.09%和25.14%/48.32%的精确率/召回率
  • 医生评估其检索文档有用性超单用大模型搜索系统

下一次就诊的ICD编码预测旨在根据已有的纵向病历记录,预判未来就诊时可能记录的标准诊断编码。该任务具有前瞻性且为多标签性质:目标记录尚未存在,可能存在多个正确编码。结构化电子健康记录(EHR)基础模型捕捉疾病复发与时间演进规律,而语言基础模型则生成灵活的诊断假设。本文提出ICD-Deepresearch,一种深度研究工作流,将上述预测基础模型与医学搜索、ICD词典相结合。由于无任何来源能揭示未来编码集合,研究通过在固定前K预算内,结合患者证据、外部临床关系与精确编码语义来评估候选转换。候选生成阶段使用SparseEHR生成EHR先验,并启动两次受控的研究扩展;独立的GPT-5直接预测提供互补候选。最终选择模块对两条路径结果进行验证、去重与联合排序,随后由独立模块撰写推理理由而不改变预测结果。最终,ICD-Deepresearch在MIMIC-III上实现平均精度/召回率为24.60%/35.09%,在MIMIC-IV上为25.14%/48.32%。医生评价其检索文档有用性分别为51%和68%,显著高于单独使用GPT-5网络搜索的22%和39%,以及医学深度研究系统的32%和41%。

原文摘要 · Abstract (English)

Next-encounter ICD forecasting predicts which standardized diagnosis codes will be documented at a future visit from the longitudinal record available beforehand. The task is prospective and multi-label: the target note does not yet exist, and several codes may be correct. Structured EHR foundation models capture recurrence and temporal progression, whereas language foundation models generate flexible diagnostic hypotheses. We introduce ICD-Deepresearch, a DeepResearch workflow that composes these predictive foundation models with medical search and ICD dictionaries. Because no source reveals the future code set, research evaluates candidate transitions by linking patient evidence, external clinical relations, and exact code semantics under a fixed top-K budget. Candidate Generation uses SparseEHR to produce an EHR Prior that initializes two bounded Research Expansion rounds; an independent GPT-5 Direct Forecast supplies complementary candidates. Final Selection validates, deduplicates, and jointly ranks both paths, after which a separate module writes rationales without changing predictions. Finally ICD-Deepresearch achieves patient-averaged precision/recall of 24.60/35.09% on MIMIC-III and 25.14/48.32% on MIMIC-IV. Physicians rate 51% and 68% of its retrieved documents useful, compared with 22% and 39% for standalone GPT-5 web search and 32% and 41% for Medical Deep Research. ICD-Deepresearch therefore improves over the registered local comparators while retrieving evidence with higher physician-rated usefulness than the standalone research systems

临床预测ICD编码多模型协同医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。