arXiv:2604.10386cs.AIcs.MA2026-04

用多智能体框架分析病历轨迹,实现多癌种早期风险预测。

TrajOnco: a multi-agent framework for temporal reasoning over longitudinal EHR for multi-cancer early detection

论文配图:TrajOnco: a multi-agent framework for temporal reasoning over longitudinal EHR for multi-cancer early detection
图 1 · 摘自论文原文
  • 采用多智能体链式架构,结合长期记忆进行时间序列推理。
  • 零样本下1年癌症预测的AUROC达0.64-0.80,媲美监督学习模型。
  • 输出可解释,适合临床研究与大规模筛查应用。

从纵向电子健康记录(EHR)中准确估算癌症风险可支持早期发现与改善诊疗,但建模复杂患者轨迹仍具挑战。我们提出TrajOnco,一种无需训练的多智能体大语言模型框架,用于可扩展的多癌种早期检测。该框架采用链式多智能体结构并具备长期记忆,对连续临床事件进行时间推理,生成患者级摘要、证据关联的推理过程及风险评分。我们在去标识化的Truveta EHR数据上评估了15种癌症类型,使用匹配的病例对照队列,预测1年内癌症诊断风险。在零样本评估中,TrajOnco的AUROC为0.64–0.80,在肺癌基准上表现堪比监督学习模型,且时间推理能力优于单智能体LLM。多智能体设计还使小型模型(如GPT-4.1-mini)实现有效推理。人工评估验证了输出的真实性。此外,其可解释的推理结果可聚合揭示与临床共识一致的人群风险模式。这些结果凸显多智能体LLM在可解释的时间推理与临床洞察生成方面的潜力,推动多癌种早期检测发展。

原文摘要 · Abstract (English)

Accurate estimation of cancer risk from longitudinal electronic health records (EHRs) could support earlier detection and improved care, but modeling such complex patient trajectories remains challenging. We present TrajOnco, a training-free, multi-agent large language model (LLM) framework designed for scalable multi-cancer early detection. Using a chain-of-agents architecture with long-term memory, TrajOnco performs temporal reasoning over sequential clinical events to generate patient-level summaries, evidence-linked rationales, and predicted risk scores. We evaluated TrajOnco on de-identified Truveta EHR data across 15 cancer types using matched case-control cohorts, predicting risk of cancer diagnosis at 1 year. In zero-shot evaluation, TrajOnco achieved AUROCs of 0.64-0.80, performing comparably to supervised machine learning in a lung cancer benchmark while demonstrating better temporal reasoning than single-agent LLMs. The multi-agent design also enabled effective temporal reasoning with smaller-capacity models such as GPT-4.1-mini. The fidelity of TrajOnco's output was validated through human evaluation. Furthermore, TrajOnco's interpretable reasoning outputs can be aggregated to reveal population-level risk patterns that align with established clinical knowledge. These findings highlight the potential of multi-agent LLMs to execute interpretable temporal reasoning over longitudinal EHRs, advancing both scalable multi-cancer early detection and clinical insight generation.

多智能体病历分析癌症预测可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。