arXiv:2505.07683cs.LGcs.AI2025-05被引 2

用大模型嵌入融合多模态癌症数据,提升预测性能。

Multimodal Cancer Modeling in the Age of Foundation Model Embeddings

  • 基于大模型嵌入,统一处理基因、影像与病理文本
  • 多模态融合显著优于单一模态,提升生存预测准确率
  • 验证文本摘要对结果影响,适合临床研究者参考

癌症基因组图谱(TCGA)通过整合基因组、临床和影像数据,成为癌症研究的重要大规模参考数据集。以往研究多针对生存预测等任务构建专用深度学习模型。当前生物医学深度学习趋势是发展基础模型(FMs),生成不依赖特定任务的特征嵌入。尤其在生物医学文本领域,基础模型发展迅速。尽管TCGA包含自由文本的病理报告,但长期未被充分利用。本文研究在多模态零样本基础模型嵌入上训练经典机器学习模型的能力。结果表明,多模态融合易于实现且具有叠加效应,显著优于单模态模型。进一步验证了纳入病理报告文本的价值,并系统评估了模型生成摘要与幻觉的影响。总体提出一种以嵌入为中心的多模态癌症建模新范式。

原文摘要 · Abstract (English)

The Cancer Genome Atlas (TCGA) has enabled novel discoveries and served as a large-scale reference dataset in cancer through its harmonized genomics, clinical, and imaging data. Numerous prior studies have developed bespoke deep learning models over TCGA for tasks such as cancer survival prediction. A modern paradigm in biomedical deep learning is the development of foundation models (FMs) to derive feature embeddings agnostic to a specific modeling task. Biomedical text especially has seen growing development of FMs. While TCGA contains free-text data as pathology reports, these have been historically underutilized. Here, we investigate the ability to train classical machine learning models over multimodal, zero-shot FM embeddings of cancer data. We demonstrate the ease and additive effect of multimodal fusion, outperforming unimodal models. Further, we show the benefit of including pathology report text and rigorously evaluate the effect of model-based text summarization and hallucination. Overall, we propose an embedding-centric approach to multimodal cancer modeling.

多模态癌症建模基础模型病理文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。