arXiv:2604.24473cs.AIcs.CL2026-04

用智能体推理分析多发性骨髓瘤长期病历,接近专家水平。

Agentic clinical reasoning over longitudinal myeloma records: a retrospective evaluation against expert consensus

  • 设计智能体系统分步推理,整合海量病历信息。
  • 准确率达79.6%,显著优于传统方法(+3.8~4.2个百分点)。
  • 复杂问题和长病历下表现更优,适合临床决策支持场景。

多发性骨髓瘤需多年持续治疗,每项决策依赖分布于数十至数百份异构临床文档中的累积病史。现有研究尚未验证大模型能否达到专家共识水平的证据整合能力。本研究回顾性评估了811例患者(2001–2026年,某三甲中心)的纵向病历,涵盖44,962份文档与1,334,677个实验室值,并在MIMIC-IV上进行外部验证。对比智能体推理系统、单次检索增强生成(RAG)、迭代RAG及全上下文输入,在48种模板下的469个患者-问题对上进行测试,参考标签由四位肿瘤科医生双盲标注并经资深血液科医师仲裁。迭代RAG与全上下文输入收敛于同一上限(75.4% vs 75.8%,p=1.00)。智能体系统达79.6%一致率(95% CI 76.4–82.8),显著超越两者(+3.8和+4.2个百分点;p=0.006、0.007)。增益随问题复杂度上升,在基于标准的综合判断中达+9.4个百分点(p=0.032);随记录长度增长,在最长记录前10%中达+13.5个百分点(n=10)。系统错误率(12.2%)与专家分歧率(13.6%)相当,但严重性相反:57.8%的系统错误具临床重要性,而专家分歧仅18.8%。智能体是唯一突破共享上限的方法,优势集中于最复杂问题与最长记录。残留错误的更高临床后果表明,需前瞻性评估其在真实诊疗中的效果后方可转化为患者获益。

原文摘要 · Abstract (English)

Multiple myeloma is managed through sequential lines of therapy over years to decades, with each decision depending on cumulative disease history distributed across dozens to hundreds of heterogeneous clinical documents. Whether LLM-based systems can synthesise this evidence at a level approaching expert agreement has not been established. A retrospective evaluation was conducted on longitudinal clinical records of 811 myeloma patients treated at a tertiary centre (2001-2026), covering 44,962 documents and 1,334,677 laboratory values, with external validation on MIMIC-IV. An agentic reasoning system was compared against single-pass retrieval-augmented generation (RAG), iterative RAG, and full-context input on 469 patient-question pairs from 48 templates at three complexity levels. Reference labels came from double annotation by four oncologists with senior haematologist adjudication. Iterative RAG and full-context input converged on a shared ceiling (75.4% vs 75.8%, p = 1.00). The agentic system reached 79.6% concordance (95% CI 76.4-82.8), exceeding both baselines (+3.8 and +4.2 pp; p = 0.006 and 0.007). Gains rose with question complexity, reaching +9.4 pp on criteria-based synthesis (p = 0.032), and with record length, reaching +13.5 pp in the top decile (n = 10). The system error rate (12.2%) was comparable to expert disagreement (13.6%), but severity was inverted: 57.8% of system errors were clinically significant versus 18.8% of expert disagreements. Agentic reasoning was the only approach to exceed the shared ceiling, with gains concentrated on the most complex questions and longest records. The greater clinical consequence of residual system errors indicates that prospective evaluation in routine care is required before these findings translate into patient benefit.

智能体临床决策多发性骨髓瘤大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。