用大模型统一建模癌症患者病程,预测更准且可解释。
TwinWeaver: An LLM-Based Foundation Model Framework for Pan-Cancer Digital Twins
- 把患者长期病历转成文本,用大模型统一预测疾病进展
- 在9.3万患者数据上,预测误差降低10%,生存风险分层准确率超基线
- 能直接用于新临床试验,无需重新训练,适合精准医疗研究者
精准肿瘤学需要预测临床事件和病程演变,但稀疏、多模态的临床时间序列建模仍是关键挑战。我们提出TwinWeaver,一个开源框架,将纵向患者病史序列化为文本,使大语言模型可统一处理事件预测与轨迹推演,并基于20种癌症类型共93,054名患者的数据构建了Genie数字孪生(GDT)。在基准测试中,GDT显著降低预测误差,中位均绝对缩放误差(MASE)为0.87,优于最强时间序列基线的0.97(p<0.001)。此外,GDT提升风险分层能力,各类任务平均一致性指数(C-index)达0.703,超过最优基线的0.662。该模型还具备分布外泛化能力,在未见过的临床试验中实现零样本匹配,微调后中位MASE降至0.75–0.88,事件预测平均C-index达0.672,超越最强基线的0.648。最后,TwinWeaver支持可解释的临床推理扩展,为长期临床建模提供可扩展、透明的基础。
原文摘要 · Abstract (English)
Precision oncology requires forecasting clinical events and trajectories, yet modeling sparse, multi-modal clinical time series remains a critical challenge. We introduce TwinWeaver, an open-source framework that serializes longitudinal patient histories into text, enabling unified event prediction as well as forecasting with large language models, and use it to build Genie Digital Twin (GDT) on 93,054 patients across 20 cancer types. In benchmarks, GDT significantly reduces forecasting error, achieving a median Mean Absolute Scaled Error (MASE) of 0.87 compared to 0.97 for the strongest time-series baseline (p<0.001). Furthermore, GDT improves risk stratification, achieving an average concordance index (C-index) of 0.703 across survival, progression, and therapy switching tasks, surpassing the best baseline of 0.662. GDT also generalizes to out-of-distribution clinical trials, matching trained baselines at zero-shot and surpassing them with fine-tuning, achieving a median MASE of 0.75-0.88 and outperforming the strongest baseline in event prediction with an average C-index of 0.672 versus 0.648. Finally, TwinWeaver enables an interpretable clinical reasoning extension, providing a scalable and transparent foundation for longitudinal clinical modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。